rescue rule sad-Sadness-S4-k2 — voice-corrected

Sadness under rescue rule S4, k=2. one hop straight across the gap, step cap lifted. 90.0 % of clips score at or below zero on this emotion and the largest gap on its normalised axis is 0.449 (WIDER than the 0.25 step cap). This rule found 19,060 chains over 40,000 tracks; the strict rule found 0 at k=3.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_sad-Sadness-S4-k2.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
These are not strict-rule trajectories. They come from a deliberately looser rule, built to recover examples on an emotion the strict rule cannot reach. What rule S4 changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop. What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression. Full explanation →
20chains converted
20segments re-voiced
0.778 → 0.824median worst-to-anchor identity cosine
92 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Sadness rising ↑identity +0.03 emotion 110 %   sad-Sadness-S4-k2 · #1

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.44. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.44.

On the corpus-wide percentile scale those become 0.43, 0.92 — a total move of +0.49.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 23 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.747 before conversion and 0.776 after — it rose by 0.028. Neighbour-to-neighbour the worst pair went 0.747 → 0.776. (The earlier render, with segment 1 left raw, scores 0.691 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.486 in the original and +0.537 after conversion — 110 % of the delta retained, i.e. the move came out slightly larger after conversion than before.

Quality. Mean predicted overall quality across the segments went 2.94 → 3.19 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 0.4404normalised 0.451 → 0.939identity cos to seg 1 0.747 → 0.776 +0.028identity cos neighbours 0.747 → 0.776d_b rescored +0.486 → +0.537d_a rescored +0.486 → +0.537d_a mined 0.440d_b mined 0.488min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_--z5fTsHDactotal 22.5schain gain +0.7 dBseam step 1.2 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normal-paced, normally alert, slightly relaxed
(some disfluency, average clarity, monologue, didactic) Und zwar stellt sich in unserem Fall die Frage, ob denn der Familienrichter, Dr. Markus Bühler,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 0.0/10; 6.2s, DE.
DE_--z5fTsHDac_W000000 · in -18.5 dBFS · gain -1.5 dB · emolia-00037
(sourness, malevolence malice, sadness · little disfluency, clear, didactic, monologue) dachte, der Vater ist der Böse und die Mutter ist die Gute, die mir auch das Essen, wahrscheinlich die Spätzle mit Soße auf den Tisch stellt. Und er hat dann ein Hass gegen seinen Vater entwickelt und eben nicht gegen die Mutter, wobei wahrscheinlich die Mutter den Vater benutzt hat.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, malevolence malice, sadness; style: didactic, monologue; average recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.3/10; 16.5s, DE.
DE_--z5fTsHDac_W000015 · in -18.3 dBFS · gain -1.7 dB · emolia-00037
Sadness rising ↑identity +0.49 emotion 96 %   sad-Sadness-S4-k2 · #2

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.79. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.79.

On the corpus-wide percentile scale those become 0.43, 0.96 — a total move of +0.53.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 16 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.308 before conversion and 0.797 after — it rose by 0.489. Neighbour-to-neighbour the worst pair went 0.308 → 0.797. (The earlier render, with segment 1 left raw, scores 0.740 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.528 in the original and +0.508 after conversion — 96 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.82 → 3.10 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 0.7896normalised 0.451 → 0.968identity cos to seg 1 0.308 → 0.797 +0.489identity cos neighbours 0.308 → 0.797d_b rescored +0.528 → +0.508d_a rescored +0.528 → +0.508d_a mined 0.790d_b mined 0.517min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-05ZmfVq_H8total 16.1schain gain +0.6 dBseam step 0.7 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, some disfluency
(anger · measured, somewhat unclear, monologue) Ja, meine Damen, meine Herren, die FDP-Bundestagsfraktion hat natürlich den Fall Nawalny im Zusammenhang auch mit den deutsch-russischen Beziehungen heute (ahem) diskutiert. (low mumble)
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger; style: monologue; average recording, no background noise; genuineness 3.3/6; vocal-burst blend 0.0/10; 11.1s, DE.
DE_-05ZmfVq_H8_W000000 · in -16.3 dBFS · gain -3.7 dB · emolia-00037
(sadness, disappointment, anger · normal-paced, average clarity, conversational, casual) Etwas, das man durchaus spürt im russischen Staatsapparat. Also, ich würde das jetzt nicht einfach so.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sadness, disappointment, anger; style: conversational, casual; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 0.0/10; 5.2s, DE.
DE_-05ZmfVq_H8_W000009 · in -12.4 dBFS · gain -7.6 dB · emolia-00037
Sadness rising ↑identity +0.02 emotion 108 %   sad-Sadness-S4-k2 · #3

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.41. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.41.

On the corpus-wide percentile scale those become 0.43, 0.91 — a total move of +0.48.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 24 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.777 before conversion and 0.796 after — it rose by 0.019. Neighbour-to-neighbour the worst pair went 0.777 → 0.796. (The earlier render, with segment 1 left raw, scores 0.756 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.483 in the original and +0.520 after conversion — 108 % of the delta retained, i.e. the move came out slightly larger after conversion than before.

Quality. Mean predicted overall quality across the segments went 2.96 → 3.09 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 0.4102normalised 0.451 → 0.937identity cos to seg 1 0.777 → 0.796 +0.019identity cos neighbours 0.777 → 0.796d_b rescored +0.483 → +0.520d_a rescored +0.483 → +0.520d_a mined 0.410d_b mined 0.486min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-6F1Ss25DJQtotal 23.6schain gain +1.1 dBseam step 0.9 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady
(no disfluency, clear, didactic, formal) Tokens sollten ja schon am 17. bzw. 18. Dezember
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, formal; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 0.0/10; 4.4s, DE.
DE_-6F1Ss25DJQ_W000000 · in -19.4 dBFS · gain -0.7 dB · emolia-00037
(emotional numbness, fear, interest · some disfluency, average clarity, monologue, casual) (low mumble) eh, dadurch soll also entsprechend die Inflation ein bisschen gestoppt werden, ist aber im Grunde alles Quatsch, denn im, gerade am Anfang ist die Inflation sicherlich höher als das, was hier reinfließen wird, geht ja auch nicht anders, ist ja jetzt erst am Anfang, (low mumble) es können ja im Prinzip nicht mehr Tokens, (low mumble) gelockt und gestaked werden als,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, fear, interest; style: monologue, casual; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 1.8/10; 19.4s, DE.
DE_-6F1Ss25DJQ_W000122 · in -20.4 dBFS · gain +0.4 dB · emolia-00037
Sadness rising ↑identity +0.64 emotion 5 %   sad-Sadness-S4-k2 · #4

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.74. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.74.

On the corpus-wide percentile scale those become 0.43, 0.95 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 40 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.023 before conversion and 0.618 after — it rose by 0.641. Neighbour-to-neighbour the worst pair went -0.023 → 0.618. (The earlier render, with segment 1 left raw, scores 0.529 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.523 in the original and +0.027 after conversion — 5 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.93 → 3.17 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 0.7427normalised 0.451 → 0.964identity cos to seg 1 -0.023 → 0.618 +0.641identity cos neighbours -0.023 → 0.618d_b rescored +0.523 → +0.027d_a rescored +0.523 → +0.027d_a mined 0.743d_b mined 0.513min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-8Xv740j-8Qtotal 39.8schain gain +0.1 dBseam step 2.0 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(fear, anger, relief · clear, monologue, casual) Ich würde sagen, da ist auf jeden Fall Bedarf der Verbesserung, dass man sich nicht schämen muss, in diesem Beruf zu arbeiten. Es lohnt sich auf jeden Fall, weil der Schmutz stirbt nicht aus.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as fear, anger, relief; style: monologue, casual; good recording, quiet background; genuineness 2.0/6; vocal-burst blend 0.0/10; 10.1s, DE.
DE_-8Xv740j-8Q_W000000 · in -17.5 dBFS · gain -2.5 dB · emolia-00037
(bitterness, jealousy and envy, pain · average clarity, monologue, authoritative) Die erste Kundin ist bereits in Beratung. Die ersten Gespräche sind angebracht und damit wird auch der Webauftritt noch weiter dann optimiert. Und wir wünschen jetzt dir auch viel Erfolg dabei, wenn du dich in der Reinigungsbranche selbstständig machen möchtest oder jemand kennst. Du kannst gerne den Link unterhalb des Videos nutzen für ein kostenfreies Erstgespräch. Wir würden dann gemeinsam als Team zunächst sicherstellen, dass der AVGS dafür beantragt wird. Was ist ein AVGS?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as bitterness, jealousy and envy, pain; style: monologue, authoritative; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 1.3/10; 30.0s, DE.
DE_-8Xv740j-8Q_W000047 · in -19.1 dBFS · gain -0.9 dB · emolia-00037
Sadness rising ↑identity +0.27 emotion REVERSED   sad-Sadness-S4-k2 · #5

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 0.49. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 0.49.

On the corpus-wide percentile scale those become 0.43, 0.92 — a total move of +0.49.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 18 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.408 before conversion and 0.678 after — it rose by 0.270. Neighbour-to-neighbour the worst pair went 0.408 → 0.678. (The earlier render, with segment 1 left raw, scores 0.507 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.492 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 3.01 → 3.16 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 0.4861normalised 0.451 → 0.943identity cos to seg 1 0.408 → 0.678 +0.270identity cos neighbours 0.408 → 0.678d_b rescored +0.492 → +0.000d_a rescored +0.492 → +0.000d_a mined 0.486d_b mined 0.492min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-8c1BTgFP3ktotal 17.7schain gain +1.3 dBseam step 2.3 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normal-paced, normally alert, slightly relaxed
(average clarity, conversational, authoritative) Und jetzt der Wochendurchblick mit Florian Streibl. Liebe Zuschauerinnen und Zuschauer, willkommen zum Wochendurchblick.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational, authoritative; good recording, quiet background; genuineness 1.6/6; vocal-burst blend 0.3/10; 6.5s, DE.
DE_-8c1BTgFP3k_W000000 · in -18.5 dBFS · gain -1.5 dB · emolia-00247
(jealousy and envy, disappointment, bitterness · clear, monologue, didactic) Doch angesichts der Aggressivität mit der Russlandspräsident Putin den Westen geradezu herausfordert, hätte ich das so klar und energisch eher vom Bundeskanzler Scholz erwartet.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, disappointment, bitterness; style: monologue, didactic; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 0.0/10; 11.4s, DE.
DE_-8c1BTgFP3k_W000013 · in -20.5 dBFS · gain +0.5 dB · emolia-00247
Sadness rising ↑identity +0.01 emotion REVERSED   sad-Sadness-S4-k2 · #6

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 0.00. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 0.00.

On the corpus-wide percentile scale those become 0.43, 0.86 — a total move of +0.43.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 12 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.751 before conversion and 0.762 after — it rose by 0.011. Neighbour-to-neighbour the worst pair went 0.751 → 0.762. (The earlier render, with segment 1 left raw, scores 0.677 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.430 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.70 → 2.90 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 0.0027normalised 0.451 → 0.901identity cos to seg 1 0.751 → 0.762 +0.011identity cos neighbours 0.751 → 0.762d_b rescored +0.430 → +0.000d_a rescored +0.430 → +0.000d_a mined 0.003d_b mined 0.450min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-ArmcowDha4total 11.3schain gain +1.8 dBseam step 1.8 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(contempt, disgust, sourness · almost no disfluency, formal, newsreading) Sehr geehrte Frau Präsidentin, als Verfasser der Stellungnahme des Auswärtigen Ausschusses will ich in erster Linie über die externen Aspekte unseres Berichtes sprechen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, disgust, sourness; style: formal, newsreading; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 0.0/10; 8.3s, DE.
DE_-ArmcowDha4_W000000 · in -17.8 dBFS · gain -2.2 dB · emolia-00224
(pride · no disfluency, formal, authoritative) Dazu gehört auch der weltweite Kampf für den Klimaschutz.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pride; style: formal, authoritative; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.0/10; 3.3s, DE.
DE_-ArmcowDha4_W000003 · in -17.0 dBFS · gain -3.0 dB · emolia-00224
Sadness rising ↑identity +0.22 emotion 106 %   sad-Sadness-S4-k2 · #7

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.38. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.38.

On the corpus-wide percentile scale those become 0.43, 0.91 — a total move of +0.48.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 21 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.514 before conversion and 0.736 after — it rose by 0.221. Neighbour-to-neighbour the worst pair went 0.514 → 0.736. (The earlier render, with segment 1 left raw, scores 0.654 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.479 in the original and +0.506 after conversion — 106 % of the delta retained, i.e. the move came out slightly larger after conversion than before.

Quality. Mean predicted overall quality across the segments went 2.94 → 3.11 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 0.3831normalised 0.451 → 0.935identity cos to seg 1 0.514 → 0.736 +0.221identity cos neighbours 0.514 → 0.736d_b rescored +0.479 → +0.506d_a rescored +0.479 → +0.506d_a mined 0.383d_b mined 0.484min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-Jk4ZBOnugMtotal 20.9schain gain -0.6 dBseam step 2.9 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(emotional numbness, helplessness, fear · some disfluency, formal, storytelling) Egal was Sie machen, Ihre VR-Brille möchte einfach nicht über Ihre Seehilfe passen.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, helplessness, fear; style: formal, storytelling; good recording, quiet background; genuineness 2.2/6; vocal-burst blend 0.0/10; 4.7s, DE.
DE_-Jk4ZBOnugM_W000000 · in -16.3 dBFS · gain -3.7 dB · emolia-00147
(contemplation, concentration, helplessness · frequent disfluency, monologue, casual) Wenn man abends nochmal kurz vorm Bett die Brille dann anzieht, ja, hat man sonst eventuell Probleme einzuschlafen und das soll dann auch noch minimiert werden dadurch, durch dieses Blutprotekt. Von daher, wie gesagt, muss man selbst entscheiden, man braucht es nicht.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, concentration, helplessness; style: monologue, casual; average recording, quiet background; genuineness 5.4/6; vocal-burst blend 0.0/10; 16.4s, DE.
DE_-Jk4ZBOnugM_W000074 · in -18.1 dBFS · gain -1.9 dB · emolia-00147
Sadness rising ↑identity +0.02 emotion 86 %   sad-Sadness-S4-k2 · #8

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.66. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.66.

On the corpus-wide percentile scale those become 0.43, 0.94 — a total move of +0.51.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 21 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.900 before conversion and 0.924 after — it rose by 0.023. Neighbour-to-neighbour the worst pair went 0.900 → 0.924. (The earlier render, with segment 1 left raw, scores 0.869 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.514 in the original and +0.441 after conversion — 86 % of the delta retained, which is most of it.

Quality. Mean predicted overall quality across the segments went 2.86 → 3.15 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 0.6636normalised 0.451 → 0.958identity cos to seg 1 0.900 → 0.924 +0.023identity cos neighbours 0.900 → 0.924d_b rescored +0.514 → +0.441d_a rescored +0.514 → +0.441d_a mined 0.664d_b mined 0.506min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-NcjdeCDcbYtotal 20.8schain gain +4.0 dBseam step 0.6 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(disappointment · brisk, dramatic, storytelling) Diese Woche gibt es unseren Beitrag zum Thema verlassene Orte etwas später wie gewohnt, aber immer noch pünktlich zum Mittwoch und daran wird sich auch in der Zukunft nichts ändern.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment; style: dramatic, storytelling; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 1.2/10; 8.9s, DE.
DE_-NcjdeCDcbY_W000000 · in -20.7 dBFS · gain +0.7 dB · emolia-00182
(disgust, fear, sourness · normal-paced, formal, monologue) im ehemaligen Eingangsbereich zur Straßenseite hatte es vor einigen Jahren stark gebrannt weswegen es aufgrund der Witterung zu einem Dach- und Deckendurchbruch kam somit konnte man gerade den vorderen Teil gar nicht mehr betreten
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, fear, sourness; style: formal, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.1/10; 12.2s, DE.
DE_-NcjdeCDcbY_W000007 · in -21.3 dBFS · gain +1.3 dB · emolia-00182
Sadness rising ↑identity +0.02 emotion 119 %   sad-Sadness-S4-k2 · #9

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.23. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.23.

On the corpus-wide percentile scale those become 0.43, 0.89 — a total move of +0.46.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 15 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.850 before conversion and 0.872 after — it rose by 0.022. Neighbour-to-neighbour the worst pair went 0.850 → 0.872. (The earlier render, with segment 1 left raw, scores 0.815 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.463 in the original and +0.550 after conversion — 119 % of the delta retained, i.e. the move came out slightly larger after conversion than before.

Quality. Mean predicted overall quality across the segments went 2.91 → 3.08 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 0.2296normalised 0.451 → 0.924identity cos to seg 1 0.850 → 0.872 +0.022identity cos neighbours 0.850 → 0.872d_b rescored +0.463 → +0.550d_a rescored +0.463 → +0.550d_a mined 0.230d_b mined 0.473min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-QLrfFrKsMItotal 14.4schain gain +0.7 dBseam step 1.6 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(emotional numbness · formal, didactic) Anfang September. Der Rübenlagerplatz der Kautfabrik ist noch leer. Die Absetzanlage außer Betrieb.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, didactic; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.2/10; 5.8s, DE.
DE_-QLrfFrKsMI_W000000 · in -19.8 dBFS · gain -0.2 dB · emolia-00147
(formal, newsreading) Jedem Rübenbauern stehen vier Prozent seiner gesamten Liefermenge an zerkleinerten Rückständen zu. Sie finden besonders als hochwertiges Viehfutterverwendung.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 8.8s, DE.
DE_-QLrfFrKsMI_W000087 · in -16.9 dBFS · gain -3.1 dB · emolia-00147
Sadness rising ↑identity +0.03 emotion REVERSED   sad-Sadness-S4-k2 · #10

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.09. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.09.

On the corpus-wide percentile scale those become 0.43, 0.88 — a total move of +0.45.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 29 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.878 before conversion and 0.908 after — it rose by 0.029. Neighbour-to-neighbour the worst pair went 0.878 → 0.908. (The earlier render, with segment 1 left raw, scores 0.870 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.449 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.85 → 3.20 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 0.0875normalised 0.451 → 0.914identity cos to seg 1 0.878 → 0.908 +0.029identity cos neighbours 0.878 → 0.908d_b rescored +0.449 → +0.000d_a rescored +0.449 → +0.000d_a mined 0.087d_b mined 0.463min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-U2hZsyyyRutotal 29.1schain gain +5.0 dBseam step 0.5 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult somewhat feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, slightly relaxed, fairly steady
(malevolence malice · measured, subdued, frequent disfluency, didactic) Heute möchte ich ein Seifenblasenpapier machen. Das heißt, ich puste Seifenblasen, bunte Seifenblasen auf ein Papier, die darauf platzen und dort ihre farbige Spur hinterlassen.
full caption & clip details
A young adult somewhat feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, slightly guarded; reads as malevolence malice; style: didactic, whispered; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.5/10; 15.6s, DE.
DE_-U2hZsyyyRu_W000000 · in -22.5 dBFS · gain +2.5 dB · emolia-00147
(emotional numbness, fatigue exhaustion, disappointment · normal-paced, normally alert, some disfluency, casual) Ich habe da eine Batterie eingelegt, aber das Ganze ist eine Fehlkonstruktion. Hier sind so Stege in der Abdeckung und es lässt sich dann jetzt hier nicht mehr drüber schieben und befestigen. Also als Kinderspielzeug ungeeignet, aber
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as emotional numbness, fatigue exhaustion, disappointment; style: casual, monologue; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 0.0/10; 13.8s, DE.
DE_-U2hZsyyyRu_W000023 · in -17.9 dBFS · gain -2.1 dB · emolia-00147
Sadness rising ↑identity +0.05 emotion REVERSED   sad-Sadness-S4-k2 · #11

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 0.05. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 0.05.

On the corpus-wide percentile scale those become 0.43, 0.87 — a total move of +0.44.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 12 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.772 before conversion and 0.821 after — it rose by 0.049. Neighbour-to-neighbour the worst pair went 0.772 → 0.821. (The earlier render, with segment 1 left raw, scores 0.731 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.443 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.71 → 3.06 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 0.0462normalised 0.451 → 0.910identity cos to seg 1 0.772 → 0.821 +0.049identity cos neighbours 0.772 → 0.821d_b rescored +0.443 → +0.000d_a rescored +0.443 → +0.000d_a mined 0.046d_b mined 0.459min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-WhG-j-UudUtotal 11.6schain gain +3.1 dBseam step 1.0 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, good recording, measured, slightly relaxed, fairly steady, moderate pitch range
(normally alert, little disfluency, average clarity, authoritative) Hello. Herzlich willkommen aus der Quantum Storm Star Wars Collection. Mein Name ist Deniz und ich möchte euch heute den
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, didactic; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 0.1/10; 6.7s, DE.
DE_-WhG-j-UudU_W000000 · in -19.3 dBFS · gain -0.7 dB · emolia-00238
(fear, malevolence malice, sourness · subdued, some disfluency, somewhat unclear, monologue) Dann können die Storytrooper sich auch hinsetzen. Naja, auf so einer langen Fahrt vielleicht auch gar nicht schlecht.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as fear, malevolence malice, sourness; style: monologue, conversational; good recording, quiet background; genuineness 2.7/6; vocal-burst blend 0.7/10; 5.1s, DE.
DE_-WhG-j-UudU_W000056 · in -20.5 dBFS · gain +0.5 dB · emolia-00238
Sadness rising ↑identity +0.04 emotion 106 %   sad-Sadness-S4-k2 · #12

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 0.26. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 0.26.

On the corpus-wide percentile scale those become 0.43, 0.90 — a total move of +0.47.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 33 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.876 before conversion and 0.915 after — it rose by 0.039. Neighbour-to-neighbour the worst pair went 0.876 → 0.915. (The earlier render, with segment 1 left raw, scores 0.881 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.466 in the original and +0.496 after conversion — 106 % of the delta retained, i.e. the move came out slightly larger after conversion than before.

Quality. Mean predicted overall quality across the segments went 2.74 → 3.18 (+0.44) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 0.2629normalised 0.451 → 0.926identity cos to seg 1 0.876 → 0.915 +0.039identity cos neighbours 0.876 → 0.915d_b rescored +0.466 → +0.496d_a rescored +0.466 → +0.496d_a mined 0.263d_b mined 0.475min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-YzjWuJuZL8total 32.8schain gain +2.3 dBseam step 0.3 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, fairly steady
(hope enthusiasm optimism, elation, interest · slightly relaxed, casual, conversational) Hallo in die Runde. Ich freue mich sehr, dass ihr dabei seid, mit mir heute wieder über Skat sprechen wollt. Da könnt ihr es auch schon sehen, mir ist demnetzt tatsächlich mal wieder eine ganz interessante Partie über den Weg gelaufen, online, wie ihr seht.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, elation, interest; style: casual, conversational; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 3.8/10; 13.1s, DE.
DE_-YzjWuJuZL8_W000000 · in -17.1 dBFS · gain -2.9 dB · emolia-00238
(relief, affection, hope enthusiasm optimism · neutral tension, casual, monologue) Nützen tut's auch nicht. Spaß am Skat hab ich trotzdem nicht verloren. War aber, das muss ich jetzt sagen, tatsächlich richtig eine Therapiesitzung hier heute für mich. Damit bin ich jetzt auch durch. Hoffe, ihr hattet euren Spaß. Hoffe, wir sehen uns bald mit weiteren interessanten Verteilungen. Wünsche euch noch einen hervorragenden Rest Wochenende und eine schöne Woche. Bis bald. Gut Blatt.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, affection, hope enthusiasm optimism; style: casual, monologue; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 3.7/10; 20.0s, DE.
DE_-YzjWuJuZL8_W000039 · in -18.2 dBFS · gain -1.8 dB · emolia-00238
Sadness rising ↑identity +0.01 emotion REVERSED   sad-Sadness-S4-k2 · #13

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.07. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.07.

On the corpus-wide percentile scale those become 0.43, 0.88 — a total move of +0.45.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 28 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.832 before conversion and 0.843 after — it rose by 0.010. Neighbour-to-neighbour the worst pair went 0.832 → 0.843. (The earlier render, with segment 1 left raw, scores 0.793 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.447 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 3.19 → 3.37 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 0.0746normalised 0.451 → 0.913identity cos to seg 1 0.832 → 0.843 +0.010identity cos neighbours 0.832 → 0.843d_b rescored +0.447 → +0.000d_a rescored +0.447 → +0.000d_a mined 0.075d_b mined 0.462min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-bnJC1eLO5Itotal 28.0schain gain +2.6 dBseam step 0.1 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, quiet background, slightly relaxed, fairly steady, some disfluency, average clarity
(normal-paced, normally alert, storytelling, monologue) Wir machen eine Geisbergrunde heute miteinander. Man kann um den ganzen Geisberg rum spazieren. Und das Video, das werden wir heute rund um den Geisberg machen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: storytelling, monologue; good recording, quiet background; genuineness 1.9/6; vocal-burst blend 1.0/10; 9.6s, DE.
DE_-bnJC1eLO5I_W000000 · in -18.2 dBFS · gain -1.8 dB · emolia-00113
(relief, contemplation, pride · measured, subdued, monologue, casual) Und wenn nachher schaust, die Dinge regeln sich oft ganz für selber, oder du warst dem Moment schon was zu tun ist, oder du warst es zu einem späteren Zeitpunkt. Und im Nachhinein sagst du immer, wa, eigentlich ist ganz gut, dass das passiert ist, weil sonst hätte das andere, viel Größere, nicht passieren können.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, contemplation, pride; style: monologue, casual; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 3.5/10; 18.7s, DE.
DE_-bnJC1eLO5I_W000029 · in -18.5 dBFS · gain -1.5 dB · emolia-00113
Sadness rising ↑identity +0.01 emotion 92 %   sad-Sadness-S4-k2 · #14

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 1.33. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.33.

On the corpus-wide percentile scale those become 0.39, 0.99 — a total move of +0.59.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 46 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.897 before conversion and 0.909 after — it rose by 0.012. Neighbour-to-neighbour the worst pair went 0.897 → 0.909. (The earlier render, with segment 1 left raw, scores 0.792 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.474 in the original and +0.434 after conversion — 92 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.82 → 3.16 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 0.3337normalised 0.451 → 0.931identity cos to seg 1 0.897 → 0.909 +0.012identity cos neighbours 0.897 → 0.909d_b rescored +0.474 → +0.434d_a rescored +0.474 → +0.434d_a mined 0.334d_b mined 0.480min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-e7x_OIHd0gtotal 45.5schain gain +2.8 dBseam step 1.7 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normally alert, slightly relaxed, fairly steady
(disgust, concentration, infatuation · brisk, clear, newsreading, formal) Am Freitag startet der CDU-Bundesparteitag in Leipzig. Der CDU-Bundesparteitag ist das höchste beschlussfassende Gremium der CDU Deutschlands. Was sich kompliziert anhört, ist eigentlich ganz einfach. Dort werden die wichtigsten inhaltlichen und personellen Entscheidungen gekriegt.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as disgust, concentration, infatuation; style: newsreading, formal; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.1/10; 15.6s, DE.
DE_-e7x_OIHd0g_W000000 · in -19.9 dBFS · gain -0.1 dB · emolia-00247
(bitterness, disappointment, triumph · normal-paced, average clarity, monologue, authoritative) Aus ganz Deutschland kommen rund 1000 Delegierte nach Leipzig.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as bitterness, disappointment, triumph; style: monologue, authoritative; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 1.9/10; 30.0s, DE.
DE_-e7x_OIHd0g_W000001 · in -18.8 dBFS · gain -1.2 dB · emolia-00247
Sadness rising ↑identity +0.07 emotion 103 %   sad-Sadness-S4-k2 · #15

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.70. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.70.

On the corpus-wide percentile scale those become 0.43, 0.95 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 13 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.779 before conversion and 0.851 after — it rose by 0.072. Neighbour-to-neighbour the worst pair went 0.779 → 0.851. (The earlier render, with segment 1 left raw, scores 0.739 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.518 in the original and +0.535 after conversion — 103 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.65 → 2.93 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 0.6968normalised 0.451 → 0.960identity cos to seg 1 0.779 → 0.851 +0.072identity cos neighbours 0.779 → 0.851d_b rescored +0.518 → +0.535d_a rescored +0.518 → +0.535d_a mined 0.697d_b mined 0.509min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-f6GmCGhSBototal 12.5schain gain +3.7 dBseam step 0.4 dBcrossfades 100 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, neutral tension
(sourness, shame, thankfulness gratitude · some disfluency, average clarity, wide pitch range, conversational) Ich will euch heute zeigen, was ich wirklich innerhalb, wenn man Mutter ist, dann hat man wenig Zeit und was ich wirklich so
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, neutral stance, neutral openness; reads as sourness, shame, thankfulness gratitude; style: conversational, casual; good recording, quiet background; genuineness 4.7/6; vocal-burst blend 2.4/10; 5.9s, DE.
DE_-f6GmCGhSBo_W000000 · in -19.7 dBFS · gain -0.3 dB · emolia-00224
(confusion, longing, sadness · frequent disfluency, somewhat unclear, moderate pitch range, conversational) (exhausted groan) euh, ich hab dann noch, in der Zeit hab ich auch noch Sami (low mumble) angemacht, von wegen, er hätte mir ja nicht geholfen, oder, (low mumble) euh,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, neutral openness; reads as confusion, longing, sadness; style: conversational, casual; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 2.9/10; 6.9s, DE.
DE_-f6GmCGhSBo_W000013 · in -26.5 dBFS · gain +6.5 dB · emolia-00224
Sadness rising ↑identity +0.79 emotion 93 %   sad-Sadness-S4-k2 · #16

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.36. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.36.

On the corpus-wide percentile scale those become 0.43, 0.91 — a total move of +0.48.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 26 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.097 before conversion and 0.885 after — it rose by 0.788. Neighbour-to-neighbour the worst pair went 0.097 → 0.885. (The earlier render, with segment 1 left raw, scores 0.781 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.476 in the original and +0.443 after conversion — 93 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 3.04 → 3.25 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 0.3562normalised 0.451 → 0.933identity cos to seg 1 0.097 → 0.885 +0.788identity cos neighbours 0.097 → 0.885d_b rescored +0.476 → +0.443d_a rescored +0.476 → +0.443d_a mined 0.356d_b mined 0.482min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-ffmw_U0FWYtotal 25.8schain gain +2.2 dBseam step 0.6 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(concentration, pride · almost no disfluency, formal, newsreading) Mehr Tempo auf dem Weg zur Klimaneutralität, das fordert die Organisation für wirtschaftliche Entwicklung und Zusammenarbeit in ihrem neuesten Umweltprüfbericht von der Bundesregierung. Ein Appell der OECD lautet deshalb mehr E-Mobilität und mehr Verkehr auf die Schiene.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, pride; style: formal, newsreading; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.0/10; 16.4s, DE.
DE_-ffmw_U0FWY_W000000 · in -16.3 dBFS · gain -3.7 dB · emolia-00147
(emotional numbness, fear, sadness · no disfluency, formal, newsreading) denn Sturzfluten wie im Ahrtal werden häufiger. Sie haben zwischen den Jahren 2000 und 2021 insgesamt einen Schaden von 71 Milliarden Euro verursacht.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, fear, sadness; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.2/10; 9.7s, DE.
DE_-ffmw_U0FWY_W000007 · in -15.5 dBFS · gain -4.5 dB · emolia-00147
Sadness rising ↑identity −0.00 emotion REVERSED   sad-Sadness-S4-k2 · #17

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.00. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.00.

On the corpus-wide percentile scale those become 0.43, 0.86 — a total move of +0.43.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 19 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.828 before conversion and 0.827 after — it fell by 0.001. Neighbour-to-neighbour the worst pair went 0.828 → 0.827. (The earlier render, with segment 1 left raw, scores 0.806 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.428 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.89 → 3.09 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 0.0005normalised 0.451 → 0.900identity cos to seg 1 0.828 → 0.827 -0.001identity cos neighbours 0.828 → 0.827d_b rescored +0.428 → +0.000d_a rescored +0.428 → +0.000d_a mined 0.001d_b mined 0.449min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-iFXnOKFbJktotal 18.6schain gain +2.4 dBseam step 1.2 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, measured, fairly steady, somewhat unclear
(normally alert, relaxed, frequent disfluency, casual) Der war sogar weg, also so hoch stand da das Wasser, (low mumble) also knapp zwei Meter hoch. (low mumble) Das war nicht so eine gute Sache.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, submissive, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 0.0/10; 8.3s, DE.
DE_-iFXnOKFbJk_W000003 · in -24.9 dBFS · gain +4.9 dB · emolia-00095
(disappointment, fatigue exhaustion, helplessness · subdued, slightly relaxed, some disfluency, monologue) Schön aktiv an diesen ganzen Themen und nutzen in der Regel LoRaWan dafür, um die Daten da zu übertragen. Genau. Wenn es keine Fragen gibt, wäre ich fertig. Vielen Dank.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, fatigue exhaustion, helplessness; style: monologue, whispered; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 0.0/10; 10.6s, DE.
DE_-iFXnOKFbJk_W000081 · in -22.6 dBFS · gain +2.6 dB · emolia-00095
Sadness rising ↑identity +0.32 emotion 99 %   sad-Sadness-S4-k2 · #18

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 0.73. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 0.73.

On the corpus-wide percentile scale those become 0.43, 0.95 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 31 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.242 before conversion and 0.560 after — it rose by 0.318. Neighbour-to-neighbour the worst pair went 0.242 → 0.560. (The earlier render, with segment 1 left raw, scores 0.500 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.522 in the original and +0.518 after conversion — 99 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.66 → 3.11 (+0.44) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 0.7324normalised 0.451 → 0.963identity cos to seg 1 0.242 → 0.560 +0.318identity cos neighbours 0.242 → 0.560d_b rescored +0.522 → +0.518d_a rescored +0.522 → +0.518d_a mined 0.732d_b mined 0.512min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-xg8lm69_K8total 31.1schain gain +3.3 dBseam step 3.3 dBcrossfades 100 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-bright, fairly smooth, average recording, energised, wide pitch range
(normal-paced, slightly relaxed, fairly steady, authoritative) Und ich rufe auf den Tagesordnungspunkt 11.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very clear, frequent disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, formal; average recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.3/10; 3.2s, DE.
DE_-xg8lm69_K8_W000000 · in -18.2 dBFS · gain -1.8 dB · emolia-00224
(sourness, interest, disappointment · brisk, neutral tension, moderately variable, dramatic) Warum hat man dieses Gefühl? Wann hat man dieses Gefühl, wenn man das Gefühl hat, dass es einen Unterschied macht, wem man sozusagen mit seiner Stimme beauftragt, für die eigenen Anliegen einzustehen? Und an dieser Stelle, glaube ich, müssen wir auch darüber unterhalten, wie können wir dieses Gefühl weiter stärken. Das ist (ahem) eine insgesamt Verantwortung, der wir nachgehen wollen, insgesamt mit der hrg-Novelle. Und insofern freue ich mich auf die Diskussion dort im Sinne von einer größeren Demokratisierung aufgrund (ahem) im Sinne einer weiteren
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as sourness, interest, disappointment; style: dramatic, ranting; average recording, some background noise; genuineness 2.2/6; vocal-burst blend 2.1/10; 28.1s, DE.
DE_-xg8lm69_K8_W000058 · in -21.1 dBFS · gain +1.1 dB · emolia-00224
Sadness rising ↑identity −0.03 emotion REVERSED   sad-Sadness-S4-k2 · #19

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.00. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.00.

On the corpus-wide percentile scale those become 0.43, 0.86 — a total move of +0.43.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 14 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.798 before conversion and 0.766 after — it fell by 0.031. Neighbour-to-neighbour the worst pair went 0.798 → 0.766. (The earlier render, with segment 1 left raw, scores 0.761 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.430 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.73 → 3.07 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 0.0033normalised 0.451 → 0.901identity cos to seg 1 0.798 → 0.766 -0.031identity cos neighbours 0.798 → 0.766d_b rescored +0.430 → +0.000d_a rescored +0.430 → +0.000d_a mined 0.003d_b mined 0.450min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-xpBDPzvHIItotal 14.1schain gain +2.2 dBseam step 1.7 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, measured, normally alert, slightly relaxed
(doubt · some disfluency, formal, didactic) Du kennst bestimmt so die Situation. Du hast schon lange irgendwas nicht mehr gegessen.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as doubt; style: formal, didactic; good recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.0/10; 5.3s, DE.
DE_-xpBDPzvHII_W000000 · in -18.9 dBFS · gain -1.1 dB · emolia-00182
(confusion, disappointment, distress · little disfluency, monologue, narration) Und dann kaufst du's, aber wie das hergestellt wurde, das interessiert uns doch nicht. Wir haben uns dafür entschieden, schmeckt oder schmeckt nicht, ja.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as confusion, disappointment, distress; style: monologue, narration; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 0.6/10; 9.0s, DE.
DE_-xpBDPzvHII_W000021 · in -19.0 dBFS · gain -1.0 dB · emolia-00182
Sadness rising ↑identity +0.00 emotion 106 %   sad-Sadness-S4-k2 · #20

This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 0.17. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 0.17.

On the corpus-wide percentile scale those become 0.43, 0.89 — a total move of +0.46.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.

What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 20 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.823 before conversion and 0.827 after — it rose by 0.005. Neighbour-to-neighbour the worst pair went 0.823 → 0.827. (The earlier render, with segment 1 left raw, scores 0.758 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.457 in the original and +0.482 after conversion — 106 % of the delta retained, i.e. the move came out slightly larger after conversion than before.

Quality. Mean predicted overall quality across the segments went 2.97 → 3.11 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 0.165normalised 0.451 → 0.920identity cos to seg 1 0.823 → 0.827 +0.005identity cos neighbours 0.823 → 0.827d_b rescored +0.457 → +0.482d_a rescored +0.457 → +0.482d_a mined 0.165d_b mined 0.469min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0068rhxjjg8total 19.2schain gain +2.5 dBseam step 0.8 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, slightly dark, balanced body, average recording, quiet background, measured, slightly relaxed, steady
(normally alert, didactic, monologue) Die zehnte Etappe der Tour de France 2022 startete in Morsin, in Haute-Savoyen.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 0.1/10; 6.9s, DE.
DE_0068rhxjjg8_W000000 · in -20.8 dBFS · gain +0.8 dB · emolia-00215
(fatigue exhaustion, disappointment, bitterness · subdued, monologue, casual) Aber nach einigen Kilometern stand ein Polizist, stoppte uns, weil die Straße noch nicht offiziell freigegeben gewesen ist. So hatten wir nochmals circa eine halbe Stunde zu warten.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, disappointment, bitterness; style: monologue, casual; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.1/10; 12.5s, DE.
DE_0068rhxjjg8_W000012 · in -23.7 dBFS · gain +3.7 dB · emolia-00215