rescue rule sad-Sadness-S1-k3 — voice-corrected

Sadness under rescue rule S1, k=3. a genuine three-step ramp, once the zero block is removed. 90.0 % of clips score at or below zero on this emotion and the largest gap on its normalised axis is 0.449 (WIDER than the 0.25 step cap). This rule found 3,162 chains over 40,000 tracks; the strict rule found 0 at k=3.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_sad-Sadness-S1-k3.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
These are not strict-rule trajectories. They come from a deliberately looser rule, built to recover examples on an emotion the strict rule cannot reach. What rule S1 changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally. What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data. Full explanation →
20chains converted
40segments re-voiced
0.688 → 0.776median worst-to-anchor identity cosine
19 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Sadness rising ↑identity +0.00 emotion 1209 %   sad-Sadness-S1-k3 · #1

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.33, 0.53, 0.70. In the first clip the scorer found only a trace of Sadness (0.33); by the last it is at 0.70.

On the corpus-wide percentile scale those become 0.90, 0.93, 0.95 — a total move of +0.04.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.23, 0.40, 0.56, a move of +0.32, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 25 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.772 before conversion and 0.775 after — it rose by 0.003. Neighbour-to-neighbour the worst pair went 0.772 → 0.834. (The earlier render, with segment 1 left raw, scores 0.738 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.044 in the original and +0.535 after conversion — 1209 % of the delta retained, i.e. the move came out slightly larger after conversion than before.

Quality. Mean predicted overall quality across the segments went 2.83 → 3.08 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.3347 → 0.5264 → 0.6968normalised 0.931 → 0.946 → 0.960identity cos to seg 1 0.772 → 0.775 +0.003identity cos neighbours 0.772 → 0.834d_b rescored +0.044 → +0.535d_a rescored +0.044 → +0.535d_a mined 0.362d_b mined 0.029min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-f6GmCGhSBototal 24.5schain gain +4.0 dBseam step 1.3 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-bright, fairly smooth, quiet background, normal-paced, neutral tension, moderately variable, normal breath
(jealousy and envy, astonishment surprise, embarrassment · normally alert, some disfluency, average clarity, conversational) Ah, wo waren wir stehen geblieben? Meine Güte, wir haben gesagt, gerade erst mal ne Runde geheult. Dann muss ich, glaub, wir waren beim Pudern, ne?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as jealousy and envy, astonishment surprise, embarrassment; style: conversational, playful; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 3.3/10; 6.1s, DE.
DE_-f6GmCGhSBo_W000002 · in -20.2 dBFS · gain +0.2 dB · emolia-00224
(shame, embarrassment, fear · energised, some disfluency, average clarity, storytelling) Ja, ich hatte Besuch gekriegt und ich war, ich schwör's euch, 10 Minuten bevor der Besuch kam, war ich nicht fertig. Ich hatte kein, kein Hemd gebügelt und ich hatte kein Makeup drauf. Ich sah aus wie Kacker und
full caption & clip details
An adolescent somewhat masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as shame, embarrassment, fear; style: storytelling, casual; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 2.1/10; 12.0s, DE.
DE_-f6GmCGhSBo_W000012 · in -24.2 dBFS · gain +4.2 dB · emolia-00224
(confusion, longing, sadness · normally alert, frequent disfluency, somewhat unclear, conversational) (exhausted groan) euh, ich hab dann noch, in der Zeit hab ich auch noch Sami (low mumble) angemacht, von wegen, er hätte mir ja nicht geholfen, oder, (low mumble) euh,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, neutral openness; reads as confusion, longing, sadness; style: conversational, casual; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 2.9/10; 6.9s, DE.
DE_-f6GmCGhSBo_W000013 · in -26.5 dBFS · gain +6.5 dB · emolia-00224
Sadness rising ↑identity +0.00 emotion 53 %   sad-Sadness-S1-k3 · #2

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.13, 0.40, 0.60. In the first clip the scorer found only a trace of Sadness (0.13); by the last it is at 0.60.

On the corpus-wide percentile scale those become 0.88, 0.91, 0.94 — a total move of +0.05.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.05 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.08, 0.29, 0.47, a move of +0.39, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 60 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.930 before conversion and 0.933 after — it rose by 0.003. Neighbour-to-neighbour the worst pair went 0.949 → 0.933. (The earlier render, with segment 1 left raw, scores 0.823 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.054 in the original and +0.028 after conversion — 53 % of the delta retained.

Quality. Mean predicted overall quality across the segments went 3.02 → 3.18 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.1267 → 0.4019 → 0.6016normalised 0.917 → 0.936 → 0.952identity cos to seg 1 0.930 → 0.933 +0.003identity cos neighbours 0.949 → 0.933d_b rescored +0.054 → +0.028d_a rescored +0.054 → +0.028d_a mined 0.475d_b mined 0.035min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-iFXnOKFbJktotal 59.7schain gain +1.7 dBseam step 1.3 dBcrossfades 150/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, slightly relaxed
(bitterness, disgust, jealousy and envy · normally alert, monologue, casual) (low mumble) Mausfallen ist auch ein lustiges Beispiel. Es gibt Mausfallen, die haben Lorawahn an Bord. Also quasi man hat das Problem, die Falle schlägt zu, man kriegt es nicht mit und die Ratte (low mumble) bleibt oder die Maus bleibt erst mal ein paar Tage in der Mausfalle drin. Das ist vielleicht nicht was, was man zu Hause haben will. Man merkt es nicht, dann (low mumble) fängt es an zu riechen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as bitterness, disgust, jealousy and envy; style: monologue, casual; average recording, quiet background; genuineness 5.6/6; vocal-burst blend 1.3/10; 17.8s, DE.
DE_-iFXnOKFbJk_W000030 · in -22.7 dBFS · gain +2.7 dB · emolia-00095
(contemplation, disappointment, confusion · subdued, casual, monologue) Ja, also, das ist, (ahem) äh, kann ich nur empfehlen, dass man da irgendwie, wenn man sich die Möglichkeit hat, in Lokal, irgendwie in so ein Hacker-Spaces, Maker-Spaces oder vielleicht ein bisschen eine Connection hat zur Stadt, dass man da irgendwie, (low mumble) äh, einfach mal fragt, geht so was. Das ist oft viel möglich, was man davor denkt, eigentlich darf man das gar nicht machen, das ist schwierig, wer erlaubt das? Ist eigentlich ganz einfach. (low mumble)
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, disappointment, confusion; style: casual, monologue; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 3.9/10; 21.3s, DE.
DE_-iFXnOKFbJk_W000066 · in -23.8 dBFS · gain +3.8 dB · emolia-00095
(doubt, disappointment, bitterness · normally alert, monologue) bewegt leicht, weil Eckerwäsche drüber fahren, weil sich der Beton irgendwie, (low mumble) äh, ausdehnt. Ist aber alles nicht der Fall. Das haben wir da noch den Brückenbeauftragten gefragt. Der Stadtohlmann hat gemeint, ne, ne, wenn ihr Schwankungen habt von drei Zentimeter, die Brücke würde auseinander brechen. Das kann gar nicht sein. Also haben wir ein bisschen nachgeforscht und sind draufgekommen. Es ist wahrscheinlich einfach die kalte und warme Luft, Luft bei zehn Grad oder bei null Grad.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, disappointment, bitterness; style: monologue; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 1.7/10; 21.1s, DE.
DE_-iFXnOKFbJk_W000069 · in -24.6 dBFS · gain +4.6 dB · emolia-00095
Sadness rising ↑identity +0.54 emotion 24 %   sad-Sadness-S1-k3 · #3

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.39, 0.56, 0.73. In the first clip the scorer found only a trace of Sadness (0.39); by the last it is at 0.73.

On the corpus-wide percentile scale those become 0.91, 0.93, 0.95 — a total move of +0.04.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.28, 0.43, 0.59, a move of +0.31, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 76 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.143 before conversion and 0.685 after — it rose by 0.541. Neighbour-to-neighbour the worst pair went 0.234 → 0.685. (The earlier render, with segment 1 left raw, scores 0.519 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.042 in the original and +0.010 after conversion — 24 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.77 → 3.26 (+0.48) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.3899 → 0.5635 → 0.7324normalised 0.935 → 0.949 → 0.963identity cos to seg 1 0.143 → 0.685 +0.541identity cos neighbours 0.234 → 0.685d_b rescored +0.042 → +0.010d_a rescored +0.042 → +0.010d_a mined 0.343d_b mined 0.028min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-xg8lm69_K8total 75.3schain gain +2.5 dBseam step 1.6 dBcrossfades 150/100 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-bright, fairly smooth, average recording, some disfluency, average clarity
(pride, intoxication altered states of consciousness, sadness · measured, subdued, slightly relaxed, monologue) Bevor ich Frau Eisenhardt von BÜNDNIS 90DIE GRÜNEN das Wort erteile, (ahem) informiere ich Sie, dass ich gebeten worden bin, den Anfang der Rede von Frau Dr. Sommer im Ältestenrat zum Thema zu machen. Ich habe offensichtlich mit dem halben Ohr etwas verpasst, aber wir reden dann im Ältestenrat darüber.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, intoxication altered states of consciousness, sadness; style: monologue; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.4/10; 17.6s, DE.
DE_-xg8lm69_K8_W000041 · in -26.6 dBFS · gain +6.6 dB · emolia-00224
(sourness, impatience and irritability, bitterness · brisk, energised, neutral tension, dramatic) Und ich finde, das müssen wir auch mitdiskutieren. Das wäre sehr spannend. Haben Sie angekündigt, (ahem) für die HAG-Novelle, also da an der Frage Rechte der Studierenden, auch noch einmal was zu machen? Da sind wir sehr gespannt drauf. Und da würde ich auch einen Schlüssel sehen. Insgesamt sehen wir aber, dass das politische Engagement teilweise sowieso an vielen Stellen (ahem) zurückgegangen ist. Also nicht nur, was die Wahlbeteiligung angeht, sondern auch die Leute, die bereit sind, sich wählen zu lassen. Und da müssen wir schon überlegen, ob das was zu tun hat mit der Verdichtung von Studiengängen, mit der Verkürzung von Studiengängen, mit der Erwachsenenzahl von Studierenden,
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, fairly guarded; reads as sourness, impatience and irritability, bitterness; style: dramatic, ranting; average recording, some background noise; genuineness 2.8/6; vocal-burst blend 4.3/10; 30.0s, DE.
DE_-xg8lm69_K8_W000053 · in -24.7 dBFS · gain +4.7 dB · emolia-00224
(sourness, interest, disappointment · brisk, energised, neutral tension, dramatic) Warum hat man dieses Gefühl? Wann hat man dieses Gefühl, wenn man das Gefühl hat, dass es einen Unterschied macht, wem man sozusagen mit seiner Stimme beauftragt, für die eigenen Anliegen einzustehen? Und an dieser Stelle, glaube ich, müssen wir auch darüber unterhalten, wie können wir dieses Gefühl weiter stärken. Das ist (ahem) eine insgesamt Verantwortung, der wir nachgehen wollen, insgesamt mit der hrg-Novelle. Und insofern freue ich mich auf die Diskussion dort im Sinne von einer größeren Demokratisierung aufgrund (ahem) im Sinne einer weiteren
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as sourness, interest, disappointment; style: dramatic, ranting; average recording, some background noise; genuineness 2.2/6; vocal-burst blend 2.1/10; 28.1s, DE.
DE_-xg8lm69_K8_W000058 · in -21.1 dBFS · gain +1.1 dB · emolia-00224
Sadness rising ↑identity −0.06 emotion REVERSED   sad-Sadness-S1-k3 · #4

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.07, 0.13, 0.40. In the first clip the scorer found only a trace of Sadness (0.07); by the last it is at 0.40.

On the corpus-wide percentile scale those become 0.88, 0.88, 0.91 — a total move of +0.04.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.02, 0.08, 0.28, a move of +0.27, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 29 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.703 before conversion and 0.640 after — it fell by 0.063. Neighbour-to-neighbour the worst pair went 0.703 → 0.640. (The earlier render, with segment 1 left raw, scores 0.594 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.036 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.77 → 3.00 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.0651 → 0.1331 → 0.3965normalised 0.912 → 0.918 → 0.936identity cos to seg 1 0.703 → 0.640 -0.063identity cos neighbours 0.703 → 0.640d_b rescored +0.036 → +0.000d_a rescored +0.036 → +0.000d_a mined 0.331d_b mined 0.024min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-xpBDPzvHIItotal 28.5schain gain +3.1 dBseam step 2.4 dBcrossfades 100/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normally alert, slightly relaxed, light breath
(helplessness · measured, fairly steady, no disfluency, formal) dass ich seit gefühlt ewigen Zeiten mal wieder Lust hatte,
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness; style: formal, monologue; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 0.4/10; 3.6s, DE.
DE_-xpBDPzvHII_W000002 · in -14.5 dBFS · gain -5.5 dB · emolia-00182
(interest, contempt, disgust · normal-paced, moderately variable, some disfluency, conversational) (low mumble) ähm, keine Ahnung, das noch irgendwo geschwungen ist, oder es ist der Plastikbescher, oder es ist die, die Plastikverpackung, die ausschaut wie so ein Naturbescher, wie auch immer. Wir entscheiden uns aufgrund unseres Gefühls. Das gefällt uns, und dann kaufen wir es. Richtig? So.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as interest, contempt, disgust; style: conversational, playful; good recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.0/10; 17.7s, DE.
DE_-xpBDPzvHII_W000006 · in -19.3 dBFS · gain -0.7 dB · emolia-00182
(pain, sadness · measured, moderately variable, some disfluency, casual) Der erste Impuls, der da ist, dein Herz, deine Seele, die sprechen immer eine ganz klare Sprache. Die können dich nicht belügen.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as pain, sadness; style: casual, conversational; good recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.0/10; 7.7s, DE.
DE_-xpBDPzvHII_W000015 · in -17.7 dBFS · gain -2.3 dB · emolia-00182
Sadness rising ↑identity +0.47 emotion 245 %   sad-Sadness-S1-k3 · #5

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.83, 1.11, 1.39. In the first clip the scorer already found some Sadness here (0.83); by the last it is at 1.39.

On the corpus-wide percentile scale those become 0.96, 0.98, 0.99 — a total move of +0.03.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.03 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.67, 0.84, 0.93, a move of +0.26, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 37 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.303 before conversion and 0.774 after — it rose by 0.470. Neighbour-to-neighbour the worst pair went 0.434 → 0.794. (The earlier render, with segment 1 left raw, scores 0.675 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.031 in the original and +0.075 after conversion — 245 % of the delta retained, i.e. the move came out slightly larger after conversion than before.

Quality. Mean predicted overall quality across the segments went 2.93 → 3.18 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.8262 → 1.1152 → 1.3867normalised 0.970 → 0.986 → 0.994identity cos to seg 1 0.303 → 0.774 +0.470identity cos neighbours 0.434 → 0.794d_b rescored +0.031 → +0.075d_a rescored +0.031 → +0.075d_a mined 0.560d_b mined 0.023min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_02FqQjJyMtytotal 35.9schain gain +4.3 dBseam step 2.0 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, balanced body, normal-paced
(anger, jealousy and envy, fear · energised, neutral tension, variable, storytelling) Ich musste verhindern, dass der Kobold meine Schwester tötet. (ahem) Er wollte mir meinen Zwilling nehmen. Wenn er denkt, dass ich Ann dadurch aufgeben werde, (exhausted groan) irt er sich gewaltig.
full caption & clip details
A middle-aged masculine voice; delivery is energised, normal-paced, neutral tension, variable; timbre is neutral-toned, neutral-bright, very rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is negative, slightly dominant, vulnerable; reads as anger, jealousy and envy, fear; style: storytelling, dramatic; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 2.6/10; 10.5s, DE.
DE_02FqQjJyMty_W000010 · in -21.1 dBFS · gain +1.1 dB · emolia-00215
(jealousy and envy, pain, distress · very low-energy, slightly relaxed, moderately variable, storytelling) Nein, es gehört ihr. Wenn ich es behalte, wäre ich so schlimm wie die Wilderer.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as jealousy and envy, pain, distress; style: storytelling, ASMR; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.8/10; 4.8s, DE.
DE_02FqQjJyMty_W000030 · in -18.4 dBFS · gain -1.6 dB · emolia-00215
(jealousy and envy, bitterness, distress · very low-energy, neutral tension, variable, casual) Rechne ich außerhalb des Schlosses jederzeit mit einem Angriff der Wilderer. Verständlich. Wir haben ja auch unseren Kampfring sabotiert. Und ihr Drachenei gestohlen. Stimmt. Du hast recht. Aber warum sind sie uns nicht gefolgt? Das sieht ihnen gar nicht ähnlich. Außer. Außer was? Außer. Sie haben es doch getan.
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, normal-paced, neutral tension, variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as jealousy and envy, bitterness, distress; style: casual, storytelling; good recording, quiet background; mildly explicit content; genuineness 2.5/6; vocal-burst blend 3.3/10; 21.0s, DE.
DE_02FqQjJyMty_W000037 · in -21.0 dBFS · gain +1.0 dB · emolia-00215
Sadness rising ↑identity +0.01 emotion REVERSED   sad-Sadness-S1-k3 · #6

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.47, 0.60, 0.79. In the first clip the scorer found only a trace of Sadness (0.47); by the last it is at 0.79.

On the corpus-wide percentile scale those become 0.92, 0.94, 0.96 — a total move of +0.04.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.35, 0.46, 0.64, a move of +0.30, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 47 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.869 before conversion and 0.876 after — it rose by 0.007. Neighbour-to-neighbour the worst pair went 0.869 → 0.902. (The earlier render, with segment 1 left raw, scores 0.764 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The emotional move did not survive. Re-scored end to end, Sadness moved +0.038 in the original and -0.002 after conversion — it changed direction. On this chain the corrected audio is not an improvement.

Quality. Mean predicted overall quality across the segments went 2.60 → 3.20 (+0.60) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.4707 → 0.5981 → 0.7925normalised 0.942 → 0.952 → 0.968identity cos to seg 1 0.869 → 0.876 +0.007identity cos neighbours 0.869 → 0.902d_b rescored +0.038 → -0.002d_a rescored +0.038 → -0.002d_a mined 0.322d_b mined 0.026min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0KZXkf63R18total 46.0schain gain +2.7 dBseam step 1.1 dBcrossfades 150/100 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, slightly relaxed, fairly steady, average clarity
(sadness · normal-paced, normally alert, some disfluency, casual) Du hast die Diagnose KPTBS, beziehungsweise leidest du unter Entwicklungstrauma und hättest du die Möglichkeit, (low mumble) ähm, Traumakonfrontation zu erleben oder durchzumachen. Stellst du dir aber die Frage, ist das was für mich?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as sadness; style: casual, monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.0/10; 15.2s, DE.
DE_0KZXkf63R18_W000000 · in -21.9 dBFS · gain +1.9 dB · emolia-00254
(shame, concentration, fatigue exhaustion · measured, very low-energy, frequent disfluency, monologue) Wenn ich unkontrolliertes (ahem) Gewaltverhalt mir selbst gegenüber habe, also mich selbst verletze, (low mumble) ähm, oder entsprechende Suchterkrankung habe, mir ja auch schädigt dauerhaft.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, audible breath; affect is mildly positive, neutral stance, neutral openness; reads as shame, concentration, fatigue exhaustion; style: monologue, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 0.0/10; 13.7s, DE.
DE_0KZXkf63R18_W000038 · in -18.2 dBFS · gain -1.8 dB · emolia-00254
(fatigue exhaustion, shame, helplessness · normal-paced, normally alert, some disfluency, monologue) doch noch versucht hatte und irgendwie nach der ersten Sitzung oder so sogar abgebrochen hat, weil das mehr kaputt machen kann, das ist ein recht intensiver Verfahren. Die Grundvoraussetzungen sind dann auch zeitintensiv. Wie gesagt, ist nicht für Komplexdramatisierte geeignet.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as fatigue exhaustion, shame, helplessness; style: monologue, casual; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.0/10; 17.4s, DE.
DE_0KZXkf63R18_W000051 · in -19.8 dBFS · gain -0.2 dB · emolia-00254
Sadness rising ↑identity +0.03 emotion REVERSED   sad-Sadness-S1-k3 · #7

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.24, 0.39, 0.64. In the first clip the scorer found only a trace of Sadness (0.24); by the last it is at 0.64.

On the corpus-wide percentile scale those become 0.89, 0.91, 0.94 — a total move of +0.05.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.05 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.17, 0.28, 0.50, a move of +0.33, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 41 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.847 before conversion and 0.872 after — it rose by 0.025. Neighbour-to-neighbour the worst pair went 0.847 → 0.872. (The earlier render, with segment 1 left raw, scores 0.798 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The emotional move did not survive. Re-scored end to end, Sadness moved +0.046 in the original and -0.021 after conversion — it changed direction. On this chain the corrected audio is not an improvement.

Quality. Mean predicted overall quality across the segments went 2.89 → 3.15 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.2441 → 0.3865 → 0.6362normalised 0.925 → 0.935 → 0.955identity cos to seg 1 0.847 → 0.872 +0.025identity cos neighbours 0.847 → 0.872d_b rescored +0.046 → -0.021d_a rescored +0.046 → -0.021d_a mined 0.392d_b mined 0.030min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0S219bZ30GQtotal 40.6schain gain +3.3 dBseam step 1.2 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(jealousy and envy, sourness, disgust · almost no disfluency, formal, monologue) Manche Menschen erfahren bei der Geburt, dass sie intersexuell sind, manche erst in der Pubertät, wenn beispielsweise die Menstruation bei einem Mädchen ausbleibt, weil es keine Eierstücke, sondern Hoden hat. Manche Menschen erfahren es im Erwachsenenalter, wenn sie medizinische Hilfe suchen, weil sie keine Kinder bekommen können. Andere erfahren es nie. Um trotzdem ein paar Zahlen zu nennen.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as jealousy and envy, sourness, disgust; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.1/10; 20.8s, DE.
DE_0S219bZ30GQ_W000016 · in -24.0 dBFS · gain +4.0 dB · emolia-00097
(contempt, fear, disgust · little disfluency, formal, monologue) Oft wird Intersexualität nämlich nur als Beweis dafür verwendet, dass es mehr als zwei Geschlechter gibt. Was vor allem Transgendermenschen zugute kommt.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, fear, disgust; style: formal, monologue; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.0/10; 8.2s, DE.
DE_0S219bZ30GQ_W000065 · in -23.4 dBFS · gain +3.4 dB · emolia-00097
(emotional numbness, pain, fear · some disfluency, formal, monologue) Das größte Problem intersexueller Menschen ist, dass viele intersexuelle Babys noch im Säuglingsalter zum weiblichen oder männlichen Geschlecht umoperiert werden. Das bedeutet nicht nur viele Schmerzen und Narben,
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, pain, fear; style: formal, monologue; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 0.0/10; 12.0s, DE.
DE_0S219bZ30GQ_W000068 · in -20.1 dBFS · gain +0.1 dB · emolia-00097
Sadness rising ↑identity −0.02 emotion 1018 %   sad-Sadness-S1-k3 · #8

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.17, 0.34, 0.61. In the first clip the scorer found only a trace of Sadness (0.17); by the last it is at 0.61.

On the corpus-wide percentile scale those become 0.89, 0.91, 0.94 — a total move of +0.05.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.05 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.11, 0.24, 0.47, a move of +0.36, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 30 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.850 before conversion and 0.826 after — it fell by 0.025. Neighbour-to-neighbour the worst pair went 0.868 → 0.846. (The earlier render, with segment 1 left raw, scores 0.744 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.049 in the original and +0.499 after conversion — 1018 % of the delta retained, i.e. the move came out slightly larger after conversion than before.

Quality. Mean predicted overall quality across the segments went 3.07 → 3.25 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.1746 → 0.345 → 0.6074normalised 0.921 → 0.932 → 0.953identity cos to seg 1 0.850 → 0.826 -0.025identity cos neighbours 0.868 → 0.846d_b rescored +0.049 → +0.499d_a rescored +0.049 → +0.499d_a mined 0.433d_b mined 0.032min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0yKDkN9ARCAtotal 29.1schain gain +3.1 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, measured, slightly relaxed, moderate pitch range
(disappointment · normally alert, fairly steady, some disfluency, monologue) Also faktisch ist es das Ziel oder die Hoffnung, dass wir da scheitern in diesem ersten Phase, weil sonst wäre es ja nicht gut, wenn man diese so einfach deanonymisieren könnte. In der zweiten Phase.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as disappointment; style: monologue, didactic; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.0/10; 12.8s, DE.
DE_0yKDkN9ARCA_W000040 · in -17.6 dBFS · gain -2.4 dB · emolia-00166
(pride, shame, longing · normally alert, fairly steady, some disfluency, monologue) Künstliche Intelligenz letzten Jahrzehnt, in den letzten zehn Jahren Vorraum geschritten ist, hat insbesondere auch im Bereich NLP hat es grosse Schritte gegeben.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, shame, longing; style: monologue, casual; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.1/10; 8.9s, DE.
DE_0yKDkN9ARCA_W000054 · in -16.5 dBFS · gain -3.5 dB · emolia-00166
(fatigue exhaustion, longing, disappointment · very low-energy, moderately variable, frequent disfluency, casual) So, nun habe wir noch, glaube ich, einige Minuten Zeit für Fragen aus dem Publikum. Und ich begebe wieder die Chance.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is mildly positive, neutral stance, neutral openness; reads as fatigue exhaustion, longing, disappointment; style: casual, monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 2.7/10; 7.8s, DE.
DE_0yKDkN9ARCA_W000084 · in -17.0 dBFS · gain -3.0 dB · emolia-00166
Sadness rising ↑identity +0.18 emotion 11 %   sad-Sadness-S1-k3 · #9

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.79, 1.08, 1.40. In the first clip the scorer already found some Sadness here (0.79); by the last it is at 1.40.

On the corpus-wide percentile scale those become 0.96, 0.98, 0.99 — a total move of +0.04.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.64, 0.83, 0.93, a move of +0.30, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 28 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.581 before conversion and 0.765 after — it rose by 0.184. Neighbour-to-neighbour the worst pair went 0.759 → 0.802. (The earlier render, with segment 1 left raw, scores 0.641 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.035 in the original and +0.004 after conversion — 11 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.55 → 2.94 (+0.40) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.7861 → 1.082 → 1.4033normalised 0.967 → 0.985 → 0.994identity cos to seg 1 0.581 → 0.765 +0.184identity cos neighbours 0.759 → 0.802d_b rescored +0.035 → +0.004d_a rescored +0.035 → +0.004d_a mined 0.617d_b mined 0.027min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0ytNrKylOrctotal 27.4schain gain +2.2 dBseam step 0.7 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an elderly somewhat feminine voice · slightly warm, smooth, balanced body, average recording, quiet background, slow, very low-energy, relaxed
(sexual lust, contemplation, fatigue exhaustion · somewhat unclear, narrow pitch range, minimal breath, whispered) Und so lasse dich tragen von den Energiewellen, welche sich bereits zu dir bewegen. Es sind die Energiewellen, welche über meinem Kanal zu dir fließen.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, smooth, balanced body; somewhat unclear, frequent disfluency, narrow pitch range, minimal breath; affect is mildly positive, submissive, neutral openness; reads as sexual lust, contemplation, fatigue exhaustion; style: whispered, ASMR; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 3.3/10; 15.1s, DE.
DE_0ytNrKylOrc_W000000 · in -16.7 dBFS · gain -3.3 dB · emolia-00215
(contemplation, affection, sadness · slurred, narrow pitch range, audible breath, ASMR) sondern, dass Mutter Erde dich begleitet und dass du wichtig bist.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, very dark, smooth, balanced body; slurred, frequent disfluency, narrow pitch range, audible breath; affect is mildly positive, submissive, vulnerable; reads as contemplation, affection, sadness; style: ASMR, whispered; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 4.4/10; 8.0s, DE.
DE_0ytNrKylOrc_W000006 · in -15.2 dBFS · gain -4.8 dB · emolia-00215
(relief, sadness, helplessness · slurred, fairly narrow pitch, light breath, whispered) All das, was dich gerade belastet jetzt von dir weg.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, slightly dark, smooth, balanced body; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is mildly negative, submissive, neutral openness; reads as relief, sadness, helplessness; style: whispered, ASMR; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 1.8/10; 4.7s, DE.
DE_0ytNrKylOrc_W000008 · in -17.4 dBFS · gain -2.5 dB · emolia-00215
Sadness rising ↑identity −0.04 emotion 60 %   sad-Sadness-S1-k3 · #10

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.51, 0.66, 0.81. In the first clip the scorer already found some Sadness here (0.51); by the last it is at 0.81.

On the corpus-wide percentile scale those become 0.92, 0.94, 0.96 — a total move of +0.04.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.38, 0.52, 0.65, a move of +0.28, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 32 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.855 before conversion and 0.816 after — it fell by 0.039. Neighbour-to-neighbour the worst pair went 0.821 → 0.742. (The earlier render, with segment 1 left raw, scores 0.787 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.036 in the original and +0.022 after conversion — 60 % of the delta retained.

Quality. Mean predicted overall quality across the segments went 2.90 → 3.14 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.5059 → 0.6577 → 0.8071normalised 0.944 → 0.957 → 0.969identity cos to seg 1 0.855 → 0.816 -0.039identity cos neighbours 0.821 → 0.742d_b rescored +0.036 → +0.022d_a rescored +0.036 → +0.022d_a mined 0.301d_b mined 0.025min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_12MSkDIQFvktotal 31.5schain gain -2.3 dBseam step 1.4 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an elderly somewhat feminine voice · slightly dark, balanced body, frequent disfluency, fairly narrow pitch
(fatigue exhaustion, contemplation, longing · slow, very low-energy, relaxed, monologue) Und uns gewarnt, ja, nicht hinausschauen und ruhig bleiben und, (ahem) äh, und wir Kinder haben das nicht verstanden und,
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, submissive, neutral openness; reads as fatigue exhaustion, contemplation, longing; style: monologue, ASMR; below-average recording, some background noise; genuineness 3.0/6; vocal-burst blend 0.9/10; 9.5s, DE.
DE_12MSkDIQFvk_W000002 · in -16.8 dBFS · gain -3.2 dB · emolia-00003
(fear, doubt, confusion · measured, subdued, slightly relaxed, whispered) Und haben sie sich nicht mitgenommen, weil sie Angst kommt von Ansteckung. Und die ist auch zurückgeblieben.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as fear, doubt, confusion; style: whispered, ASMR; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.2/10; 5.5s, DE.
DE_12MSkDIQFvk_W000012 · in -16.4 dBFS · gain -3.6 dB · emolia-00003
(fatigue exhaustion, sadness, contemplation · slow, very low-energy, relaxed, whispered) Wo sie unsre verhaftet haben, da haben sie so viel Leute verhaftet, dass ein Pferdestahl in der Nähe von Eisenkarpel voll war. Und die haben sie dann mit Lastwegen abgeführt. Nach Klagenfurt und dann in verschiedene Kassette.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, submissive, neutral openness; reads as fatigue exhaustion, sadness, contemplation; style: whispered, ASMR; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.0/10; 17.0s, DE.
DE_12MSkDIQFvk_W000013 · in -20.6 dBFS · gain +0.7 dB · emolia-00003
Sadness rising ↑identity +0.43 emotion 1070 %   sad-Sadness-S1-k3 · #11

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.13, 0.26, 0.56. In the first clip the scorer found only a trace of Sadness (0.13); by the last it is at 0.56.

On the corpus-wide percentile scale those become 0.88, 0.90, 0.93 — a total move of +0.05.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.05 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.08, 0.18, 0.42, a move of +0.35, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 20 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.037 before conversion and 0.391 after — it rose by 0.428. Neighbour-to-neighbour the worst pair went 0.109 → 0.441. (The earlier render, with segment 1 left raw, scores 0.169 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.047 in the original and +0.507 after conversion — 1070 % of the delta retained, i.e. the move came out slightly larger after conversion than before.

Quality. Mean predicted overall quality across the segments went 2.67 → 2.85 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.1294 → 0.261 → 0.5571normalised 0.918 → 0.926 → 0.949identity cos to seg 1 -0.037 → 0.391 +0.428identity cos neighbours 0.109 → 0.441d_b rescored +0.047 → +0.507d_a rescored +0.047 → +0.507d_a mined 0.428d_b mined 0.031min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1Pd71w8hAhItotal 18.8schain gain +3.1 dBseam step 0.8 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-bright, fairly smooth, quiet background, normally alert, some disfluency
(embarrassment, fatigue exhaustion, infatuation · normal-paced, neutral tension, moderately variable, casual) Ich drück mich immer falsch auf, aus, in dieser, also, filterraum, ich muss auch mal die richtige Sprache lernen.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, normal breath; affect is positive, slightly submissive, slightly vulnerable; reads as embarrassment, fatigue exhaustion, infatuation; style: casual, conversational; below-average recording, quiet background; genuineness 5.3/6; vocal-burst blend 4.9/10; 6.3s, DE.
DE_1Pd71w8hAhI_W000000 · in -21.4 dBFS · gain +1.4 dB · emolia-00077
(confusion, doubt, fear · fast, neutral tension, moderately variable, casual) Wenn was im Inkubator kommt, im Labor oder so, das ist doch irgendwie so wie beim Brutkasten, wo da im Reagenz, (ahem) in diesem Tellerchen, dass er schneller wächst und so.
full caption & clip details
A young adult feminine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as confusion, doubt, fear; style: casual, playful; below-average recording, quiet background; genuineness 5.1/6; vocal-burst blend 5.8/10; 9.4s, DE.
DE_1Pd71w8hAhI_W000129 · in -16.9 dBFS · gain -3.0 dB · emolia-00077
(astonishment surprise, awe, longing · normal-paced, slightly relaxed, fairly steady, casual) Und dann noch diese Erhöhungen, die da immer kommen, mein Gott, ja.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as astonishment surprise, awe, longing; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 3.2/10; 3.5s, DE.
DE_1Pd71w8hAhI_W000187 · in -19.4 dBFS · gain -0.6 dB · emolia-00077
Sadness rising ↑identity +0.25 emotion 1289 %   sad-Sadness-S1-k3 · #12

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.36, 0.45, 0.68. In the first clip the scorer found only a trace of Sadness (0.36); by the last it is at 0.68.

On the corpus-wide percentile scale those become 0.91, 0.92, 0.95 — a total move of +0.04.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.25, 0.33, 0.54, a move of +0.28, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 42 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.258 before conversion and 0.506 after — it rose by 0.249. Neighbour-to-neighbour the worst pair went 0.274 → 0.506. (The earlier render, with segment 1 left raw, scores 0.473 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.039 in the original and +0.506 after conversion — 1289 % of the delta retained, i.e. the move came out slightly larger after conversion than before.

Quality. Mean predicted overall quality across the segments went 2.91 → 3.21 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.3606 → 0.4478 → 0.6772normalised 0.933 → 0.940 → 0.959identity cos to seg 1 0.258 → 0.506 +0.249identity cos neighbours 0.274 → 0.506d_b rescored +0.039 → +0.506d_a rescored +0.039 → +0.506d_a mined 0.317d_b mined 0.025min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1ZMIiMMqy9utotal 40.8schain gain +2.5 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, balanced body, average recording, normal breath
(contempt, impatience and irritability, sourness · normal-paced, energised, slightly relaxed, dramatic) Wer das nicht tut oder sogar aus religiösem Fanatismus heraus Straftaten vorbereitet oder dazu in Hasspredigten anstiftet, muss unser Land schnellstmöglich verlassen, meine Damen und Herren.
full caption & clip details
A middle-aged masculine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, rough, balanced body; very clear, almost no disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as contempt, impatience and irritability, sourness; style: dramatic, monologue; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 0.2/10; 12.1s, DE.
DE_1ZMIiMMqy9u_W000034 · in -22.1 dBFS · gain +2.1 dB · emolia-00123
(anger, bitterness, concentration · normal-paced, energised, neutral tension, authoritative) Und es garantiert das Recht auf freie Religionsausübung nicht nur Menschen einer Glaubensrichtung, sondern allen. Und nicht nur, weil wir dem verpflichtet sind, sondern weil das unsere feste Überzeugung ist, hat Rassismus und haben alle anderen Formen von Ausgrenzung und Diskriminierung in Hessen keinen Platz.
full caption & clip details
A middle-aged masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as anger, bitterness, concentration; style: authoritative, monologue; average recording, some background noise; genuineness 1.0/6; vocal-burst blend 0.9/10; 20.8s, DE.
DE_1ZMIiMMqy9u_W000083 · in -21.7 dBFS · gain +1.7 dB · emolia-00123
(emotional numbness, disgust, disappointment · measured, normally alert, slightly relaxed, didactic) sind auch Alltagsdiskriminierungen ausgesetzt, beim Einkaufen, im öffentlichen Nahverkehr, im Vorbeigehen auf der Straße.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as emotional numbness, disgust, disappointment; style: didactic, monologue; average recording, no background noise; genuineness 2.0/6; vocal-burst blend 0.3/10; 8.4s, DE.
DE_1ZMIiMMqy9u_W000090 · in -22.5 dBFS · gain +2.5 dB · emolia-00123
Sadness rising ↑identity +0.07 emotion 14 %   sad-Sadness-S1-k3 · #13

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.11, 0.26, 0.51. In the first clip the scorer found only a trace of Sadness (0.11); by the last it is at 0.51.

On the corpus-wide percentile scale those become 0.88, 0.90, 0.92 — a total move of +0.04.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.06, 0.18, 0.38, a move of +0.32, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 56 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.794 before conversion and 0.860 after — it rose by 0.066. Neighbour-to-neighbour the worst pair went 0.794 → 0.860. (The earlier render, with segment 1 left raw, scores 0.809 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.043 in the original and +0.006 after conversion — 14 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 3.05 → 3.26 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.1084 → 0.2629 → 0.5098normalised 0.916 → 0.926 → 0.945identity cos to seg 1 0.794 → 0.860 +0.066identity cos neighbours 0.794 → 0.860d_b rescored +0.043 → +0.006d_a rescored +0.043 → +0.006d_a mined 0.401d_b mined 0.029min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1f41NBH2oCEtotal 55.3schain gain +4.0 dBseam step 0.2 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · neutral-toned, fairly smooth, balanced body, good recording, quiet background, wide pitch range, light breath
(contempt, anger, malevolence malice · normal-paced, normally alert, slightly relaxed, formal) Und zwar von Anfang an. Du hast keine Ahnung, was Safe Files sind? Keine Panik! In diesem Video erkläre ich dir alles Schritt für Schritt. Zuerst einmal die Basics, was Safe Files überhaupt sind und wo du diese finden kannst. Und dann natürlich auch, wie du sie in dein Spiel bekommst. Los geht's!
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as contempt, anger, malevolence malice; style: formal, playful; good recording, quiet background; genuineness 0.2/6; vocal-burst blend 0.3/10; 19.6s, DE.
DE_1f41NBH2oCE_W000001 · in -19.8 dBFS · gain -0.2 dB · emolia-00077
(disappointment, confusion, doubt · brisk, normally alert, neutral tension, casual) find ich absolut großartig. Wir werden uns jetzt natürlich nicht alle Häuser im Detail anschauen, aber ich finde, was man hier sieht, sieht schon alles wirklich sehr, sehr gut aus. Aber bevor ich das Video beende, möchte ich unbedingt noch bei der Familie Pancakes vorbeischauen und einfach schauen, was da so passiert ist.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as disappointment, confusion, doubt; style: casual, dramatic; good recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.7/10; 18.1s, DE.
DE_1f41NBH2oCE_W000043 · in -21.2 dBFS · gain +1.2 dB · emolia-00077
(interest, hope enthusiasm optimism, jealousy and envy · brisk, energised, slightly relaxed, playful) Das ist nämlich das Großartige an Save Files, dass so viele Geschichten hier verpackt sind. Also, dass die Sims auch eigene Geschichten haben und eigene Beziehungen und andere Stammbäume. Das macht eben alles um so viel spannender. Und hier sind wir. Ah, das soll ne Doppelhaushälfte sein.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as interest, hope enthusiasm optimism, jealousy and envy; style: playful, conversational; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.9/10; 18.0s, DE.
DE_1f41NBH2oCE_W000044 · in -20.4 dBFS · gain +0.3 dB · emolia-00077
Sadness rising ↑identity +0.03 emotion 45 %   sad-Sadness-S1-k3 · #14

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.67, 0.75, 1.12. In the first clip the scorer already found some Sadness here (0.67); by the last it is at 1.12.

On the corpus-wide percentile scale those become 0.94, 0.95, 0.98 — a total move of +0.04.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.53, 0.61, 0.85, a move of +0.32, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 44 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.626 before conversion and 0.656 after — it rose by 0.030. Neighbour-to-neighbour the worst pair went 0.726 → 0.772. (The earlier render, with segment 1 left raw, scores 0.657 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.039 in the original and +0.017 after conversion — 45 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.82 → 3.02 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.6689 → 0.7515 → 1.1221normalised 0.958 → 0.965 → 0.986identity cos to seg 1 0.626 → 0.656 +0.030identity cos neighbours 0.726 → 0.772d_b rescored +0.039 → +0.017d_a rescored +0.039 → +0.017d_a mined 0.453d_b mined 0.028min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1gZLllrev9ctotal 43.3schain gain +3.7 dBseam step 1.0 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, balanced body, average recording, quiet background, slightly relaxed
(helplessness, distress, fear · measured, normally alert, fairly steady, monologue) Es kann lediglich in einem sehr starken Konflikt mit meiner eigenen subjektiven Moral stehen.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness, distress, fear; style: monologue, storytelling; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 0.0/10; 6.1s, DE.
DE_1gZLllrev9c_W000005 · in -18.8 dBFS · gain -1.2 dB · emolia-00259
(malevolence malice, contemplation, helplessness · slow, subdued, steady, whispered) Welches mir persönlich am wichtigsten ist und meine gesamte Moral zusammenfasst. Die Menschheit soll existieren.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; clear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, contemplation, helplessness; style: whispered, monologue; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 0.6/10; 9.9s, DE.
DE_1gZLllrev9c_W000010 · in -21.2 dBFS · gain +1.2 dB · emolia-00259
(bitterness, helplessness, sourness · measured, subdued, steady, monologue) Dies lässt sich als evolutionärer Prozess betrachten, in dem die Moral überlebt, welche ihren Trägern die besten Voraussetzungen zum Überleben gibt und jede Moral durch beispielsweise Missverständnisse langsam mutiert. Da die Existenz ihrer Träger das Axiom meiner Moral ist, ist sie zwingend immer am besten dafür geeignet und wird sich, ob nun durch mich oder nicht, irgendwann durchsetzen.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, fairly guarded; reads as bitterness, helplessness, sourness; style: monologue, narration; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 0.1/10; 27.7s, DE.
DE_1gZLllrev9c_W000017 · in -19.2 dBFS · gain -0.8 dB · emolia-00259
Sadness rising ↑identity +0.13 emotion 12 %   sad-Sadness-S1-k3 · #15

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.28, 0.54, 0.79. In the first clip the scorer found only a trace of Sadness (0.28); by the last it is at 0.79.

On the corpus-wide percentile scale those become 0.90, 0.93, 0.96 — a total move of +0.06.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.06 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.19, 0.41, 0.64, a move of +0.45, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 23 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.668 before conversion and 0.796 after — it rose by 0.128. Neighbour-to-neighbour the worst pair went 0.668 → 0.795. (The earlier render, with segment 1 left raw, scores 0.709 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.061 in the original and +0.008 after conversion — 12 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.89 → 3.06 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.2754 → 0.5391 → 0.7876normalised 0.927 → 0.947 → 0.968identity cos to seg 1 0.668 → 0.796 +0.128identity cos neighbours 0.668 → 0.795d_b rescored +0.061 → +0.008d_a rescored +0.061 → +0.008d_a mined 0.512d_b mined 0.040min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1jA6oeQrc1Atotal 22.7schain gain +3.1 dBseam step 1.3 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normal-paced, normally alert, slightly relaxed
(confusion, relief, impatience and irritability · moderately variable, wide pitch range, conversational, casual) Okay, das nervt mich jetzt. Dann gehen wir wieder nach Hause und sagt mir bitte, was ich falsch gemacht hab. Ich hab nämlich ehrlich gesagt absolut keine Ahnung.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as confusion, relief, impatience and irritability; style: conversational, casual; good recording, no background noise; genuineness 3.3/6; vocal-burst blend 0.0/10; 7.0s, DE.
DE_1jA6oeQrc1A_W000037 · in -21.4 dBFS · gain +1.4 dB · emolia-00167
(sadness, contempt, disappointment · fairly steady, moderate pitch range, conversational, casual) Emmas Partner hat schlecht auf die Verkündung von Emmas Schwangerschaft reagiert. Es ist furchtbar. Emma ist so glücklich über das Baby, aber ihr Partner will das Baby nicht.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as sadness, contempt, disappointment; style: conversational, casual; good recording, quiet background; genuineness 4.5/6; vocal-burst blend 1.3/10; 9.1s, DE.
DE_1jA6oeQrc1A_W000074 · in -23.4 dBFS · gain +3.4 dB · emolia-00167
(confusion, sadness, disappointment · fairly steady, moderate pitch range, casual, dramatic) Wir werden ihm sagen, dass er natürlich der Vater vom Kind ist, falls er das nicht schon verstanden hat. Wer weiß, vielleicht ist er ja doch nicht der Hellste.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, sadness, disappointment; style: casual, dramatic; good recording, quiet background; genuineness 4.0/6; vocal-burst blend 0.0/10; 7.1s, DE.
DE_1jA6oeQrc1A_W000077 · in -20.5 dBFS · gain +0.5 dB · emolia-00167
Sadness rising ↑identity +0.06 emotion REVERSED   sad-Sadness-S1-k3 · #16

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.09, 0.39, 0.62. In the first clip the scorer found only a trace of Sadness (0.09); by the last it is at 0.62.

On the corpus-wide percentile scale those become 0.88, 0.91, 0.94 — a total move of +0.06.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.06 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.04, 0.28, 0.48, a move of +0.44, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 43 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.817 before conversion and 0.874 after — it rose by 0.058. Neighbour-to-neighbour the worst pair went 0.817 → 0.874. (The earlier render, with segment 1 left raw, scores 0.753 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The emotional move did not survive. Re-scored end to end, Sadness moved +0.061 in the original and -0.441 after conversion — it changed direction. On this chain the corrected audio is not an improvement.

Quality. Mean predicted overall quality across the segments went 2.78 → 3.21 (+0.43) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.0892 → 0.395 → 0.6211normalised 0.914 → 0.936 → 0.954identity cos to seg 1 0.817 → 0.874 +0.058identity cos neighbours 0.817 → 0.874d_b rescored +0.061 → -0.441d_a rescored +0.061 → -0.441d_a mined 0.532d_b mined 0.039min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_28I6DLcapyctotal 42.7schain gain +3.0 dBseam step 0.4 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, fairly steady, moderate pitch range
(sourness, fatigue exhaustion, bitterness · neutral tension, some disfluency, somewhat unclear, casual) Und er merkt gar nicht, dass er, dass er 50 Stunden mehr an diesem Teil gearbeitet hat, als er an sich musste, um, ich sag mal, den Gewinn auch nachher dem, im Gesamtpreis unterzubringen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sourness, fatigue exhaustion, bitterness; style: casual, monologue; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 0.2/10; 11.2s, DE.
DE_28I6DLcapyc_W000011 · in -17.1 dBFS · gain -2.9 dB · emolia-00097
(contempt, malevolence malice, anger · slightly relaxed, some disfluency, average clarity, monologue) Das der Mangel an dieser gesellschaftlichen Anerkennung, den wir hier feststellen, hat natürlich etwas damit zu tun, dass wir sehen, dass die Wissenswirtschaft und Gesellschaft, so wie wir sie soziologisch beschreiben können, sehr stark mit einer Tendenz der Verkopfung von Bildung zu tun hat.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contempt, malevolence malice, anger; style: monologue, formal; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 0.0/10; 20.3s, DE.
DE_28I6DLcapyc_W000064 · in -17.2 dBFS · gain -2.8 dB · emolia-00097
(anger, distress, malevolence malice · slightly relaxed, almost no disfluency, clear, monologue) Und hat dann auch diesen terziären Bereich der Berufsbildung herausgestellt und anerkannt. Aber da bedarf es auch noch einige Kraftanstrengungen, um dieses Ungleichgewicht wieder ins Lot zu bringen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as anger, distress, malevolence malice; style: monologue, formal; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.0/10; 11.6s, DE.
DE_28I6DLcapyc_W000081 · in -17.8 dBFS · gain -2.2 dB · emolia-00097
Sadness rising ↑identity +0.62 emotion REVERSED   sad-Sadness-S1-k3 · #17

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.72, 0.82, 1.20. In the first clip the scorer already found some Sadness here (0.72); by the last it is at 1.20.

On the corpus-wide percentile scale those become 0.95, 0.96, 0.99 — a total move of +0.04.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.58, 0.67, 0.88, a move of +0.30, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 17 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.152 before conversion and 0.777 after — it rose by 0.625. Neighbour-to-neighbour the worst pair went 0.187 → 0.777. (The earlier render, with segment 1 left raw, scores 0.677 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.036 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.95 → 3.07 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.7212 → 0.8237 → 1.2021normalised 0.962 → 0.970 → 0.989identity cos to seg 1 0.152 → 0.777 +0.625identity cos neighbours 0.187 → 0.777d_b rescored +0.036 → +0.000d_a rescored +0.036 → +0.000d_a mined 0.481d_b mined 0.027min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_2HgcLfDY9l4total 15.9schain gain +0.6 dBseam step 0.2 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, some disfluency
(helplessness, distress, sadness · measured, clear, monologue, formal) Das erste Mal zum Start. Wie seid ihr überhaupt? Oder wie ist eure Firma auf das gekommen?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness, distress, sadness; style: monologue, formal; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 0.0/10; 4.7s, DE.
DE_2HgcLfDY9l4_W000010 · in -19.9 dBFS · gain -0.1 dB · emolia-00196
(longing, affection, distress · normal-paced, average clarity, monologue, casual) Und entsprechend dort sind einfach die anderen Faktoren sehr wichtig. Wie halt, sie müssen sprachlich halt auch ein gewisses Niveau haben.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing, affection, distress; style: monologue, casual; good recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.3/10; 6.7s, DE.
DE_2HgcLfDY9l4_W000059 · in -19.2 dBFS · gain -0.8 dB · emolia-00196
(sadness, helplessness, longing · normal-paced, average clarity, didactic, formal) Das familiäre, und, (low mumble) ähm, ich finde, die Arbeit hat total gut reingepasst, weil wir suchen ja.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as sadness, helplessness, longing; style: didactic, formal; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 1.9/10; 4.9s, DE.
DE_2HgcLfDY9l4_W000065 · in -17.3 dBFS · gain -2.7 dB · emolia-00196
Sadness rising ↑identity +0.08 emotion REVERSED   sad-Sadness-S1-k3 · #18

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.44, 0.58, 0.78. In the first clip the scorer found only a trace of Sadness (0.44); by the last it is at 0.78.

On the corpus-wide percentile scale those become 0.92, 0.93, 0.96 — a total move of +0.04.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.32, 0.45, 0.63, a move of +0.31, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 34 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.692 before conversion and 0.770 after — it rose by 0.078. Neighbour-to-neighbour the worst pair went 0.720 → 0.753. (The earlier render, with segment 1 left raw, scores 0.592 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Sadness moved +0.041 in the original and -0.526 after conversion — it changed direction. On this chain the corrected audio is not an improvement.

Quality. Mean predicted overall quality across the segments went 2.73 → 3.05 (+0.32) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.4424 → 0.584 → 0.7832normalised 0.939 → 0.951 → 0.967identity cos to seg 1 0.692 → 0.770 +0.078identity cos neighbours 0.720 → 0.753d_b rescored +0.041 → -0.526d_a rescored +0.041 → -0.526d_a mined 0.341d_b mined 0.028min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_2R_mAsRw2RItotal 33.5schain gain +1.3 dBseam step 1.3 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, some disfluency
(contentment, longing, fatigue exhaustion · slightly relaxed, fairly steady, light breath, casual) Ich dusche jeden zweiten Tag, also immer abends. Ich bin so ein Abendduscher und den Tag, den ich nicht dusche, mit Haare waschen, da wasche ich mich dann immer nach dem Sport, vor dem Anziehen einfach dann nochmal am Körper ab, damit ich einfach auch frisch in den Tag starten kann.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contentment, longing, fatigue exhaustion; style: casual, monologue; good recording, quiet background; genuineness 2.4/6; vocal-burst blend 0.3/10; 16.4s, DE.
DE_2R_mAsRw2RI_W000008 · in -15.0 dBFS · gain -5.0 dB · emolia-00258
(fatigue exhaustion, contentment, pleasure ecstasy · neutral tension, moderately variable, normal breath, casual) Das mach ich für mich, um mich wohl zu fühlen, um mich frischer zu fühlen und auch um frischer auszusehen. Es ist nichts weltbewegendes, es ist innerhalb von 10 Minuten ist mein tägliches Make-up erledigt. Und, (low mumble) uhm,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as fatigue exhaustion, contentment, pleasure ecstasy; style: casual, monologue; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 0.0/10; 13.3s, DE.
DE_2R_mAsRw2RI_W000059 · in -15.0 dBFS · gain -5.0 dB · emolia-00258
(helplessness, distress, longing · slightly relaxed, fairly steady, light breath, casual) Kommt da nichts ran, also ich höhne auch nicht trocken nach dem Duschen, das lasse ich immer lufttrocknen.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness, distress, longing; style: casual, conversational; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 1.0/10; 4.2s, DE.
DE_2R_mAsRw2RI_W000075 · in -17.5 dBFS · gain -2.5 dB · emolia-00258
Sadness rising ↑identity −0.05 emotion REVERSED   sad-Sadness-S1-k3 · #19

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.54, 0.78, 0.95. In the first clip the scorer already found some Sadness here (0.54); by the last it is at 0.95.

On the corpus-wide percentile scale those become 0.93, 0.96, 0.97 — a total move of +0.04.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.41, 0.63, 0.76, a move of +0.35, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 16 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.683 before conversion and 0.635 after — it fell by 0.048. Neighbour-to-neighbour the worst pair went 0.774 → 0.686. (The earlier render, with segment 1 left raw, scores 0.566 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Sadness moved +0.045 in the original and -0.549 after conversion — it changed direction. On this chain the corrected audio is not an improvement.

Quality. Mean predicted overall quality across the segments went 2.87 → 2.94 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.5366 → 0.7773 → 0.9531normalised 0.947 → 0.967 → 0.979identity cos to seg 1 0.683 → 0.635 -0.048identity cos neighbours 0.774 → 0.686d_b rescored +0.045 → -0.549d_a rescored +0.045 → -0.549d_a mined 0.416d_b mined 0.032min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_2UhFJjC9wtMtotal 15.7schain gain +1.9 dBseam step 1.4 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, quiet background
(confusion, fatigue exhaustion, embarrassment · slow, very low-energy, relaxed, casual) warum hat der, ich hab jetzt nicht gerechnet, dass er so viel, (low mumble) äh, ja, so reagiert.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is neutral-toned, very dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, submissive, vulnerable; reads as confusion, fatigue exhaustion, embarrassment; style: casual, conversational; below-average recording, quiet background; genuineness 3.7/6; vocal-burst blend 0.2/10; 7.1s, DE.
DE_2UhFJjC9wtM_W000016 · in -17.2 dBFS · gain -2.8 dB · emolia-00216
(thankfulness gratitude, confusion, helplessness · measured, very low-energy, relaxed, casual) Oh, beinahe unter das Auto verloren. Aber ich glaube, dass der ist, das ist das wichtigste. Und der nicht.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as thankfulness gratitude, confusion, helplessness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.3/6; vocal-burst blend 1.5/10; 5.2s, DE.
DE_2UhFJjC9wtM_W000048 · in -17.1 dBFS · gain -2.9 dB · emolia-00216
(fatigue exhaustion, confusion, helplessness · measured, subdued, slightly relaxed, casual) Oh, jetzt fehlt der, jetzt fehlt der (ahem) vierte Gang. Na toll.
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, slightly thin; slurred, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; reads as fatigue exhaustion, confusion, helplessness; style: casual, conversational; average recording, quiet background; genuineness 5.4/6; vocal-burst blend 1.3/10; 3.8s, DE.
DE_2UhFJjC9wtM_W000054 · in -17.8 dBFS · gain -2.2 dB · emolia-00216
Sadness rising ↑identity +0.70 emotion 1189 %   sad-Sadness-S1-k3 · #20

This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.41, 0.61, 0.74. In the first clip the scorer found only a trace of Sadness (0.41); by the last it is at 0.74.

On the corpus-wide percentile scale those become 0.91, 0.94, 0.95 — a total move of +0.04.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.29, 0.47, 0.60, a move of +0.31, which is usable again.

What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.

What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 38 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.077 before conversion and 0.781 after — it rose by 0.705. Neighbour-to-neighbour the worst pair went -0.037 → 0.563. (The earlier render, with segment 1 left raw, scores 0.666 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.041 in the original and +0.488 after conversion — 1189 % of the delta retained, i.e. the move came out slightly larger after conversion than before.

Quality. Mean predicted overall quality across the segments went 2.90 → 3.26 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.4072 → 0.6074 → 0.7451normalised 0.937 → 0.953 → 0.964identity cos to seg 1 0.077 → 0.781 +0.705identity cos neighbours -0.037 → 0.563d_b rescored +0.041 → +0.488d_a rescored +0.041 → +0.488d_a mined 0.338d_b mined 0.027min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_2_pDO2tImiMtotal 37.6schain gain +2.4 dBseam step 1.4 dBcrossfades 150/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, balanced body, quiet background, fairly steady
(contemplation, impatience and irritability, jealousy and envy · measured, subdued, slightly relaxed, conversational) Wenn ich an meine (low mumble) letzte oder an meine (ahem) Fikarstelle denke in Billerbeck, das hat mich immer schwer beeindruckt. Zu den Weihnachtstagen nahmen die Menschen und nehmen es auch heute noch Weihwasser mit aus der Kirche.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contemplation, impatience and irritability, jealousy and envy; style: conversational, casual; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 1.2/10; 15.6s, DE.
DE_2_pDO2tImiM_W000052 · in -18.4 dBFS · gain -1.6 dB · emolia-00196
(shame, embarrassment, sadness · normal-paced, normally alert, slightly relaxed, conversational) Ich bin jetzt gemeint, ich mit meinem Tier dabei. (low mumble) Und das ist, glaube ich, auch das Erfolgsgeheimnis dieser Sache. Also das ist der Hauptteil, würde ich sagen.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as shame, embarrassment, sadness; style: conversational, casual; good recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.0/10; 7.8s, DE.
DE_2_pDO2tImiM_W000086 · in -17.2 dBFS · gain -2.8 dB · emolia-00196
(helplessness, fear, confusion · measured, very low-energy, relaxed, monologue) Das geht in keinem Ritus und war auch von uns nicht vorgesehen. Aber dann standen die Kinder mit ihren leuchtenden Augen und (low mumble) ihren Tieren, die sie warm hielten, vor mir. Und dann kann ich nicht sagen, die sehe ich nicht. Das geht nicht.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness, fear, confusion; style: monologue, whispered; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 0.4/10; 14.6s, DE.
DE_2_pDO2tImiM_W000087 · in -18.0 dBFS · gain -2.0 dB · emolia-00196