rescue rule sad-Sadness-S3-k3 — voice-corrected

Sadness under rescue rule S3, k=3. the same, with one unconstrained clip in between. 90.0 % of clips score at or below zero on this emotion and the largest gap on its normalised axis is 0.449 (WIDER than the 0.25 step cap). This rule found 5,806 chains over 40,000 tracks; the strict rule found 0 at k=3.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_sad-Sadness-S3-k3.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
These are not strict-rule trajectories. They come from a deliberately looser rule, built to recover examples on an emotion the strict rule cannot reach. What rule S3 changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'. What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained. Full explanation →
20chains converted
40segments re-voiced
0.749 → 0.744median worst-to-anchor identity cosine
97 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Sadness rising ↑identity −0.11 emotion 96 %   sad-Sadness-S3-k3 · #1

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 0.60, 1.09. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.09.

On the corpus-wide percentile scale those become 0.43, 0.94, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 19 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.748 before conversion and 0.635 after — it fell by 0.113. Neighbour-to-neighbour the worst pair went 0.744 → 0.642. (The earlier render, with segment 1 left raw, scores 0.637 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.551 in the original and +0.530 after conversion — 96 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.85 → 3.05 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.0 → 0.6035 → 1.0869normalised 0.451 → 0.952 → 0.985identity cos to seg 1 0.748 → 0.635 -0.113identity cos neighbours 0.744 → 0.642d_b rescored +0.551 → +0.530d_a rescored +0.551 → +0.530d_a mined 1.087d_b mined 0.534min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-WhG-j-UudUtotal 17.9schain gain +2.3 dBseam step 0.1 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, measured, slightly relaxed
(normally alert, fairly steady, little disfluency, authoritative) Hello. Herzlich willkommen aus der Quantum Storm Star Wars Collection. Mein Name ist Deniz und ich möchte euch heute den
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, didactic; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 0.1/10; 6.7s, DE.
DE_-WhG-j-UudU_W000000 · in -19.3 dBFS · gain -0.7 dB · emolia-00238
(fatigue exhaustion, longing, helplessness · normally alert, fairly steady, some disfluency, monologue) auf der anderen Stirn Seite sehen wir noch einmal die Dachkanone und noch einen weiteren Einblick in das Fahrzeug.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as fatigue exhaustion, longing, helplessness; style: monologue, didactic; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 0.4/10; 8.5s, DE.
DE_-WhG-j-UudU_W000016 · in -18.1 dBFS · gain -1.9 dB · emolia-00238
(helplessness, sadness, longing · subdued, steady, no disfluency, formal) Gummibändern festgemacht, die wir von.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; reads as helplessness, sadness, longing; style: formal, monologue; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 2.0/10; 3.1s, DE.
DE_-WhG-j-UudU_W000024 · in -16.3 dBFS · gain -3.7 dB · emolia-00238
Sadness rising ↑identity +0.10 emotion 97 %   sad-Sadness-S3-k3 · #2

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 0.57, 1.19. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.19.

On the corpus-wide percentile scale those become 0.43, 0.93, 0.99 — a total move of +0.56.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 60 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.304 before conversion and 0.407 after — it rose by 0.103. Neighbour-to-neighbour the worst pair went 0.059 → 0.407. (The earlier render, with segment 1 left raw, scores 0.380 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.556 in the original and +0.539 after conversion — 97 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.78 → 3.20 (+0.42) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.0 → 0.5713 → 1.1855normalised 0.451 → 0.950 → 0.989identity cos to seg 1 0.304 → 0.407 +0.103identity cos neighbours 0.059 → 0.407d_b rescored +0.556 → +0.539d_a rescored +0.556 → +0.539d_a mined 1.185d_b mined 0.538min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-xg8lm69_K8total 59.0schain gain +3.2 dBseam step 2.9 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-bright, fairly smooth, balanced body, average recording, light breath
(normal-paced, energised, slightly relaxed, authoritative) Und ich rufe auf den Tagesordnungspunkt 11.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very clear, frequent disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, formal; average recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.3/10; 3.2s, DE.
DE_-xg8lm69_K8_W000000 · in -18.2 dBFS · gain -1.8 dB · emolia-00224
(interest, disappointment, bitterness · normal-paced, normally alert, slightly relaxed, newsreading) Legitimation und auch die Akzeptanz in der Bevölkerung, und zwar auf allen demokratischen Ebenen, wird entsprechend gestärkt. Eine Stärkung von Wahlbeteiligung hat es aber trotz allem bereits jetzt schon an der Philips-Universität in Marburg und, wie wir gehört haben, auch in Gießen gegeben. Im Frühjahr beispielsweise 2020 ist beschlossen worden, die Wahlen zum Senat und zu den Fachbereichsräten nicht als Urnenwahl durchzuführen, sondern
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as interest, disappointment, bitterness; style: newsreading, monologue; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 0.6/10; 26.1s, DE.
DE_-xg8lm69_K8_W000036 · in -23.5 dBFS · gain +3.5 dB · emolia-00224
(disappointment, distress, impatience and irritability · brisk, energised, neutral tension, monologue) Und da würde ich sagen, liegt es nicht nur daran, dass das Wahlprozedere jetzt irgendwie kompliziert wäre. Ich meine, die Studierenden bekommen die Unterlagen geschickt. Es gibt tagelang Zeit, seine Stimme abzugeben. Ich denke, wir müssen auch überlegen, ob es was damit zu tun hat, (ahem) dass die Entscheidungskompetenzen in der Verfasst Studierendenschaft nicht gerade sehr ausgeprägt sind, um es vorsichtig zu sagen. Also, den ganzen Autonomieprozess, den es gab, sind ja Kompetenzen vom Ministerium an die Hochschulen (ahem) (ahem) verlagert worden, aber eben dort vor allem an die Präsidien und an die Hochschulräte. Und ich finde, man muss vielleicht
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as disappointment, distress, impatience and irritability; style: monologue, dramatic; average recording, some background noise; genuineness 2.1/6; vocal-burst blend 2.7/10; 30.0s, DE.
DE_-xg8lm69_K8_W000052 · in -24.0 dBFS · gain +4.0 dB · emolia-00224
Sadness rising ↑identity +0.49 emotion 99 %   sad-Sadness-S3-k3 · #3

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.67, 1.16. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.16.

On the corpus-wide percentile scale those become 0.43, 0.94, 0.99 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 26 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.173 before conversion and 0.663 after — it rose by 0.490. Neighbour-to-neighbour the worst pair went 0.256 → 0.684. (The earlier render, with segment 1 left raw, scores 0.487 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.555 in the original and +0.549 after conversion — 99 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.73 → 3.04 (+0.32) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0 → 0.6694 → 1.1602normalised 0.451 → 0.958 → 0.988identity cos to seg 1 0.173 → 0.663 +0.490identity cos neighbours 0.256 → 0.684d_b rescored +0.555 → +0.549d_a rescored +0.555 → +0.549d_a mined 1.160d_b mined 0.537min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_02FqQjJyMtytotal 25.0schain gain +3.6 dBseam step 0.5 dBcrossfades 100/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, some disfluency
(measured, normally alert, slightly relaxed, casual) (low mumble) euh, wie machen wir uns hier, wir sprechen jetzt Sebastian, Sebastian Kruppter. Das ist unsere nächste Quest. Das machen wir einfach jetzt in der Aufnahme weiter.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 1.6/10; 8.6s, DE.
DE_02FqQjJyMty_W000003 · in -22.4 dBFS · gain +2.4 dB · emolia-00215
(fear, distress, anger · measured, very low-energy, neutral tension, narration) dieses Relikt, schwarze Magie oder nicht, wird Ann retten, den Fluch umkehren. Ich werde Ann nicht verlieren. Ich schicke Ann das Wappen, dann weiß sie, dass wir uns treffen müssen.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, neutral tension, variable; timbre is neutral-toned, neutral-bright, rough, full; clear, some disfluency, wide pitch range, light breath; affect is negative, neutral stance, neutral openness; reads as fear, distress, anger; style: narration, storytelling; good recording, quiet background; genuineness 1.2/6; vocal-burst blend 3.4/10; 11.9s, DE.
DE_02FqQjJyMty_W000011 · in -17.7 dBFS · gain -2.3 dB · emolia-00215
(confusion, pain, helplessness · normal-paced, energised, neutral tension, storytelling) Ach, nichts, nur ein Gedanke. Jetzt bin ich fest entschlossen, die Macht dieses Relikts aufzuspüren.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as confusion, pain, helplessness; style: storytelling, conversational; good recording, no background noise; mildly explicit content; genuineness 1.6/6; vocal-burst blend 2.8/10; 5.0s, DE.
DE_02FqQjJyMty_W000012 · in -18.4 dBFS · gain -1.6 dB · emolia-00215
Sadness rising ↑identity +0.02 emotion 99 %   sad-Sadness-S3-k3 · #4

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.02, 1.02. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.02.

On the corpus-wide percentile scale those become 0.43, 0.87, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 30 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.767 before conversion and 0.787 after — it rose by 0.020. Neighbour-to-neighbour the worst pair went 0.610 → 0.618. (The earlier render, with segment 1 left raw, scores 0.622 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.548 in the original and +0.543 after conversion — 99 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 3.06 → 3.17 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0 → 0.018 → 1.0244normalised 0.451 → 0.906 → 0.982identity cos to seg 1 0.767 → 0.787 +0.020identity cos neighbours 0.610 → 0.618d_b rescored +0.548 → +0.543d_a rescored +0.548 → +0.543d_a mined 1.024d_b mined 0.531min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0FuG3tlx0yQtotal 28.9schain gain +4.2 dBseam step 1.3 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, quiet background, normally alert, average clarity, light breath
(normal-paced, slightly relaxed, fairly steady, conversational) auch wenn's nicht stattfindet, ein bisschen Lagerstellen, man darf doch hier aufkommen. Lass uns doch jetzt gemeinsam kurz ein Zelt aufstellen. Los geht's!
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: conversational, authoritative; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.0/10; 6.7s, DE.
DE_0FuG3tlx0yQ_W000000 · in -24.3 dBFS · gain +4.3 dB · emolia-00097
(fatigue exhaustion, doubt, contemplation · measured, neutral tension, fairly steady, monologue) Man schaut auf die Uhr und denkt, oh, schon wieder fünf Stunden vorbei. Manchmal schleicht sie aber auch langsam wie eine Schnecke. Man weiß gar nicht, wann die Zeit endlich vorbei ist. Ich möchte euch heute mitnehmen zu einem Bibeltext, zu einem Menschen, zum Mose, der bestimmt noch keine Uhr hatte.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, doubt, contemplation; style: monologue, casual; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.4/10; 19.2s, DE.
DE_0FuG3tlx0yQ_W000002 · in -20.6 dBFS · gain +0.6 dB · emolia-00097
(sadness, triumph, helplessness · normal-paced, slightly relaxed, moderately variable, authoritative) Doch selbst noch die besten Jahre sind voll Kummer und Schmerz.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sadness, triumph, helplessness; style: authoritative, dramatic; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.7/10; 3.3s, DE.
DE_0FuG3tlx0yQ_W000006 · in -18.1 dBFS · gain -1.9 dB · emolia-00097
Sadness rising ↑identity +0.04 emotion 93 %   sad-Sadness-S3-k3 · #5

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.56, 1.09. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.09.

On the corpus-wide percentile scale those become 0.43, 0.93, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 40 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.750 before conversion and 0.791 after — it rose by 0.042. Neighbour-to-neighbour the worst pair went 0.766 → 0.791. (The earlier render, with segment 1 left raw, scores 0.806 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.552 in the original and +0.516 after conversion — 93 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 3.08 → 3.23 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0 → 0.5645 → 1.0928normalised 0.451 → 0.949 → 0.985identity cos to seg 1 0.750 → 0.791 +0.042identity cos neighbours 0.766 → 0.791d_b rescored +0.552 → +0.516d_a rescored +0.552 → +0.516d_a mined 1.093d_b mined 0.534min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0IWgiqWfPNAtotal 39.4schain gain +2.5 dBseam step 1.3 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, average clarity
(moderately variable, some disfluency, wide pitch range, conversational) Hey Leute, willkommen zurück zu einem neuen Video hier auf Popel mit Zucker. Wie immer fangen wir an mit einem extra Schluck.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: conversational, playful; good recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.4/10; 5.6s, DE.
DE_0IWgiqWfPNA_W000000 · in -17.9 dBFS · gain -2.1 dB · emolia-00195
(jealousy and envy, bitterness, astonishment surprise · fairly steady, frequent disfluency, moderate pitch range, casual) Sein, sein Chef ist weg. Der, der ihn befragt im Rollstuhl, also der andere. Das ist der kleine Junge, glaub ich, der da rumhockt. (surprised gasp) Und der ist halt auch traumatisiert und der hat jetzt rausgefunden, dass Woods, der Alte, halt Mason erschossen hat und gar nicht, 嗯, nanandis, nenandis? Mit Namen hab ich's gar nicht, aber der auf jeden Fall.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as jealousy and envy, bitterness, astonishment surprise; style: casual, monologue; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 0.0/10; 21.9s, DE.
DE_0IWgiqWfPNA_W000050 · in -19.6 dBFS · gain -0.5 dB · emolia-00195
(fatigue exhaustion, pain, sadness · fairly steady, some disfluency, moderate pitch range, monologue) zu leicht, glaube ich, meiner Meinung nach. Hätte man doppelt überlegen müssen, aber klar, man will die Rache, man will ihn endlich loswerden, nach all der Zeit, und dann laut man halt auch so eine, so eine Täuschung mal eher, aber das ist halt,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, pain, sadness; style: monologue, casual; average recording, no background noise; genuineness 3.1/6; vocal-burst blend 1.1/10; 12.3s, DE.
DE_0IWgiqWfPNA_W000059 · in -20.7 dBFS · gain +0.7 dB · emolia-00195
Sadness rising ↑identity −0.02 emotion 100 %   sad-Sadness-S3-k3 · #6

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 0.36, 1.01. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.01.

On the corpus-wide percentile scale those become 0.43, 0.91, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 28 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.921 before conversion and 0.897 after — it fell by 0.024. Neighbour-to-neighbour the worst pair went 0.911 → 0.897. (The earlier render, with segment 1 left raw, scores 0.869 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.547 in the original and +0.545 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.70 → 3.04 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.0 → 0.3594 → 1.0098normalised 0.451 → 0.933 → 0.982identity cos to seg 1 0.921 → 0.897 -0.024identity cos neighbours 0.911 → 0.897d_b rescored +0.547 → +0.545d_a rescored +0.547 → +0.545d_a mined 1.010d_b mined 0.530min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0S219bZ30GQtotal 27.4schain gain +3.1 dBseam step 1.2 dBcrossfades 100/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(steady, no disfluency, newsreading, formal) Zusammen wollen wir uns verschiedene Begriffe aus dem queeren und feministischen Spektrum anschauen, Hintergründe zu den Themen beleuchten und Fragen klären.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.2/10; 7.3s, DE.
DE_0S219bZ30GQ_W000000 · in -22.0 dBFS · gain +2.0 dB · emolia-00097
(emotional numbness, contempt, anger · fairly steady, little disfluency, didactic, formal) Das wird von intersexuellen Menschen als Begriff schon eher akzeptiert. Im Deutschen gibt es sonst noch die Begriffe zwischengeschlechtlich oder intergeschlechtlich. Sind das nicht Zwitter oder Hermaphroditen?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, contempt, anger; style: didactic, formal; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.4/10; 12.7s, DE.
DE_0S219bZ30GQ_W000009 · in -22.5 dBFS · gain +2.5 dB · emolia-00097
(contempt, sadness, distress · fairly steady, no disfluency, formal, newsreading) Die Begriffe Zwitter und Hermaphrodit sind für die meisten intersexuellen Menschen sehr beleidigend und verletzend. Du solltest sie also nicht benutzen.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, sadness, distress; style: formal, newsreading; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.2/10; 7.7s, DE.
DE_0S219bZ30GQ_W000012 · in -21.3 dBFS · gain +1.3 dB · emolia-00097
Sadness rising ↑identity −0.00 emotion REVERSED   sad-Sadness-S3-k3 · #7

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.45, 1.30. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.30.

On the corpus-wide percentile scale those become 0.43, 0.92, 0.99 — a total move of +0.56.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 31 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.777 before conversion and 0.773 after — it fell by 0.004. Neighbour-to-neighbour the worst pair went 0.777 → 0.773. (The earlier render, with segment 1 left raw, scores 0.690 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Sadness moved +0.560 in the original and -0.019 after conversion — it changed direction. On this chain the corrected audio is not an improvement.

Quality. Mean predicted overall quality across the segments went 2.76 → 3.03 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0 → 0.4541 → 1.2949normalised 0.451 → 0.940 → 0.992identity cos to seg 1 0.777 → 0.773 -0.004identity cos neighbours 0.777 → 0.773d_b rescored +0.560 → -0.019d_a rescored +0.560 → -0.019d_a mined 1.295d_b mined 0.541min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0fB4TCjiFIktotal 30.5schain gain +1.8 dBseam step 1.0 dBcrossfades 150/100 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(concentration, hope enthusiasm optimism, contemplation · monologue, formal) Ihr Lieben, (low mumble) für uns Grüne war schon immer klar, die Zukunft, die wir wollen, die müssen wir auch selber erschaffen. Und auch wenn die Folgen der Corona-Krise uns noch lange beschäftigen werden, wenn nicht jetzt alle Probleme, alle Krisen und alle Herausforderungen gleichzeitig angeht, der verspielt die Zukunft der zukünftigen Generation.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, hope enthusiasm optimism, contemplation; style: monologue, formal; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.0/10; 18.3s, DE.
DE_0fB4TCjiFIk_W000000 · in -15.9 dBFS · gain -4.1 dB · emolia-00077
(bitterness, anger, disappointment · monologue, authoritative) die Hitzeperioden, die Trockenperioden, die erzwingen ein genauso (ahem) radikales und vernünftiges Handeln.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as bitterness, anger, disappointment; style: monologue, authoritative; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.0/10; 6.3s, DE.
DE_0fB4TCjiFIk_W000001 · in -16.2 dBFS · gain -3.8 dB · emolia-00077
(pride, helplessness, sadness · didactic, monologue) Und wir sind keine kleine Partei mehr. Wir sind wahnsinnig gewachsen. Als Nina und ich als Landesvorsitzende gewählt wurden.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, helplessness, sadness; style: didactic, monologue; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.0/10; 6.3s, DE.
DE_0fB4TCjiFIk_W000010 · in -14.9 dBFS · gain -5.1 dB · emolia-00077
Sadness rising ↑identity −0.08 emotion 99 %   sad-Sadness-S3-k3 · #8

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.96, 1.67. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.67.

On the corpus-wide percentile scale those become 0.43, 0.97, 1.00 — a total move of +0.57.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 15 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.867 before conversion and 0.787 after — it fell by 0.081. Neighbour-to-neighbour the worst pair went 0.867 → 0.787. (The earlier render, with segment 1 left raw, scores 0.751 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.567 in the original and +0.562 after conversion — 99 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.78 → 2.99 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0 → 0.9629 → 1.668normalised 0.451 → 0.979 → 0.998identity cos to seg 1 0.867 → 0.787 -0.081identity cos neighbours 0.867 → 0.787d_b rescored +0.567 → +0.562d_a rescored +0.567 → +0.562d_a mined 1.668d_b mined 0.546min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_12MSkDIQFvktotal 14.6schain gain -2.4 dBseam step 2.0 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged somewhat feminine voice · neutral-toned, slightly dark, fairly smooth, balanced body, average recording, very low-energy, relaxed, fairly steady
(fatigue exhaustion, confusion, doubt · measured, average clarity, fairly narrow pitch, ASMR) Nachbarin auch mitnehmen und weiter, weiter viertn.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, submissive, neutral openness; reads as fatigue exhaustion, confusion, doubt; style: ASMR, whispered; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 0.2/10; 4.5s, DE.
DE_12MSkDIQFvk_W000006 · in -16.6 dBFS · gain -3.4 dB · emolia-00003
(confusion, doubt, longing · measured, somewhat unclear, moderate pitch range, ASMR) zu uns kummen, is, i weiss nit, wer's wäre, wenn wir da drinnen.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, submissive, slightly vulnerable; reads as confusion, doubt, longing; style: ASMR, whispered; average recording, no background noise; genuineness 3.4/6; vocal-burst blend 4.3/10; 4.1s, DE.
DE_12MSkDIQFvk_W000008 · in -17.2 dBFS · gain -2.8 dB · emolia-00003
(sadness, pain, helplessness · slow, somewhat unclear, moderate pitch range, ASMR) Kleine Kinder alle angeblie, alleine geblieben waren. Der älteste war, war sechs Jahre.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, submissive, neutral openness; reads as sadness, pain, helplessness; style: ASMR, whispered; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 0.8/10; 6.3s, DE.
DE_12MSkDIQFvk_W000009 · in -17.0 dBFS · gain -3.0 dB · emolia-00003
Sadness rising ↑identity −0.08 emotion REVERSED   sad-Sadness-S3-k3 · #9

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 0.64, 1.05. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.05.

On the corpus-wide percentile scale those become 0.43, 0.94, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 20 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.908 before conversion and 0.831 after — it fell by 0.077. Neighbour-to-neighbour the worst pair went 0.923 → 0.811. (The earlier render, with segment 1 left raw, scores 0.810 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.549 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 3.00 → 3.15 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.0 → 0.6392 → 1.0508normalised 0.451 → 0.955 → 0.983identity cos to seg 1 0.908 → 0.831 -0.077identity cos neighbours 0.923 → 0.811d_b rescored +0.549 → +0.000d_a rescored +0.549 → +0.000d_a mined 1.051d_b mined 0.532min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_163DYBR_hRgtotal 19.6schain gain +1.4 dBseam step 1.4 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(steady, formal, monologue) Startschuss für DSDS. Nicht mehr lange, und Florian Silbereisen-Fans kommen auch bei RTL auf ihre Kosten
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.9s, DE.
DE_163DYBR_hRg_W000000 · in -16.4 dBFS · gain -3.6 dB · emolia-00161
(pain, fear, sadness · steady, formal, authoritative) Florie ist aufgeregt. Florian Silbereisen? Schlimmes Lampenfieber. Je älter ich werde, desto schlimmer wird es.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as pain, fear, sadness; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 8.0s, DE.
DE_163DYBR_hRg_W000002 · in -15.0 dBFS · gain -5.0 dB · emolia-00161
(confusion, helplessness, sadness · fairly steady, storytelling, formal) Ich sitze hier und weiß genauso wenig wie ihr, was heute passiert.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as confusion, helplessness, sadness; style: storytelling, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.5/10; 4.1s, DE.
DE_163DYBR_hRg_W000011 · in -16.4 dBFS · gain -3.6 dB · emolia-00161
Sadness rising ↑identity −0.07 emotion 92 %   sad-Sadness-S3-k3 · #10

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.48, 1.05. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.05.

On the corpus-wide percentile scale those become 0.43, 0.92, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 30 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.860 before conversion and 0.795 after — it fell by 0.066. Neighbour-to-neighbour the worst pair went 0.869 → 0.823. (The earlier render, with segment 1 left raw, scores 0.735 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.550 in the original and +0.505 after conversion — 92 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.96 → 3.10 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0 → 0.4832 → 1.0527normalised 0.451 → 0.943 → 0.984identity cos to seg 1 0.860 → 0.795 -0.066identity cos neighbours 0.869 → 0.823d_b rescored +0.550 → +0.505d_a rescored +0.550 → +0.505d_a mined 1.053d_b mined 0.532min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1A7Boronwiktotal 29.4schain gain +2.1 dBseam step 1.0 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an elderly somewhat feminine voice · slightly warm, slightly dark, smooth, balanced body, no background noise, slow, very low-energy, relaxed
(contemplation, awe, sexual lust · little disfluency, clear, ASMR, monologue) was wir in der Dualität so sehr benötigen. Und wenn du bereit bist,
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, slightly dark, smooth, balanced body; clear, little disfluency, narrow pitch range, audible breath; affect is mildly positive, submissive, neutral openness; reads as contemplation, awe, sexual lust; style: ASMR, monologue; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 4.2/10; 7.5s, DE.
DE_1A7Boronwik_W000003 · in -18.7 dBFS · gain -1.3 dB · emolia-00254
(thankfulness gratitude, contentment, relief · frequent disfluency, somewhat unclear, ASMR, whispered) Du hast später im Gebet auch noch die Möglichkeit, deine eigenen Worte zu sprechen. Und so beginnen wir jetzt.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, slightly dark, smooth, balanced body; somewhat unclear, frequent disfluency, narrow pitch range, audible breath; affect is mildly positive, submissive, neutral openness; reads as thankfulness gratitude, contentment, relief; style: ASMR, whispered; average recording, no background noise; genuineness 1.9/6; vocal-burst blend 3.1/10; 10.9s, DE.
DE_1A7Boronwik_W000004 · in -21.5 dBFS · gain +1.5 dB · emolia-00254
(contentment, contemplation, longing · little disfluency, somewhat unclear, ASMR, whispered) Um dir zuerst zu danken dafür, dass ich in dieser aktuellen Zeit inkarnieren durfte.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, slightly dark, smooth, balanced body; somewhat unclear, little disfluency, narrow pitch range, audible breath; affect is mildly positive, submissive, neutral openness; reads as contentment, contemplation, longing; style: ASMR, whispered; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 2.0/10; 11.4s, DE.
DE_1A7Boronwik_W000005 · in -18.9 dBFS · gain -1.1 dB · emolia-00254
Sadness rising ↑identity +0.34 emotion 85 %   sad-Sadness-S3-k3 · #11

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.80, 1.20. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.20.

On the corpus-wide percentile scale those become 0.43, 0.96, 0.99 — a total move of +0.56.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 14 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.086 before conversion and 0.422 after — it rose by 0.336. Neighbour-to-neighbour the worst pair went 0.201 → 0.427. (The earlier render, with segment 1 left raw, scores 0.396 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.557 in the original and +0.476 after conversion — 85 % of the delta retained, which is most of it.

Quality. Mean predicted overall quality across the segments went 2.58 → 2.76 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0 → 0.8022 → 1.1973normalised 0.451 → 0.969 → 0.989identity cos to seg 1 0.086 → 0.422 +0.336identity cos neighbours 0.201 → 0.427d_b rescored +0.557 → +0.476d_a rescored +0.557 → +0.476d_a mined 1.197d_b mined 0.538min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1Pd71w8hAhItotal 13.1schain gain +0.9 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · neutral-toned, neutral-bright, fairly smooth, average recording, quiet background
(impatience and irritability, anger · normal-paced, normally alert, slightly relaxed, conversational) Ach so. Auf der Suche nach dem a priori.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as impatience and irritability, anger; style: conversational, casual; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 1.6/10; 3.1s, DE.
DE_1Pd71w8hAhI_W000001 · in -15.7 dBFS · gain -4.3 dB · emolia-00077
(helplessness, sadness, emotional numbness · normal-paced, normally alert, slightly relaxed, casual) war anders ausgerichtet, mussten wir grad mal helfen, wie deine.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness, sadness, emotional numbness; style: casual; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 0.4/10; 3.9s, DE.
DE_1Pd71w8hAhI_W000010 · in -18.1 dBFS · gain -1.9 dB · emolia-00077
(helplessness, fatigue exhaustion, distress · measured, very low-energy, neutral tension, conversational) (low mumble) euhm, also, ich hab so das Gefühl, ich lass mir da auch Zeit, ja, das drängt mich niemand. (breathy giggle)
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, vulnerable; reads as helplessness, fatigue exhaustion, distress; style: conversational, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 1.9/10; 6.5s, DE.
DE_1Pd71w8hAhI_W000014 · in -17.7 dBFS · gain -2.3 dB · emolia-00077
Sadness rising ↑identity +0.22 emotion 95 %   sad-Sadness-S3-k3 · #12

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 0.97, 1.15. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.15.

On the corpus-wide percentile scale those become 0.43, 0.97, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 38 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.278 before conversion and 0.496 after — it rose by 0.219. Neighbour-to-neighbour the worst pair went 0.278 → 0.502. (The earlier render, with segment 1 left raw, scores 0.431 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.555 in the original and +0.525 after conversion — 95 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.92 → 3.18 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.0 → 0.9678 → 1.1514normalised 0.451 → 0.979 → 0.987identity cos to seg 1 0.278 → 0.496 +0.219identity cos neighbours 0.278 → 0.502d_b rescored +0.555 → +0.525d_a rescored +0.555 → +0.525d_a mined 1.151d_b mined 0.536min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1ZMIiMMqy9utotal 37.2schain gain +2.8 dBseam step 1.9 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-bright, fairly smooth, balanced body, average recording
(normal-paced, normally alert, slightly relaxed, authoritative) Antrag der Fraktion Bündnis 90, die GRÜNEN.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, no audible breath; affect is neutral, dominant, slightly guarded; no dominant emotion; style: authoritative, formal; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.6/10; 3.1s, DE.
DE_1ZMIiMMqy9u_W000000 · in -20.9 dBFS · gain +0.9 dB · emolia-00123
(anger, bitterness, disappointment · brisk, highly aroused, neutral tension, ranting) Erleben derzeit in unserem Land permanente Grenzüberschreitungen, Verbreitungen von Hass, Diskriminierung und offenem Rassismus. Im Internet, bei Facebook, bei Twitter oder auch auf öffentlichen Veranstaltungen werden Kübel von Dreck über Menschen ausgeschüttet, die eine andere Hautfarbe haben.
full caption & clip details
An adult masculine voice; delivery is highly aroused, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; very clear, almost no disfluency, wide pitch range, normal breath; affect is elated, dominant, fairly guarded; reads as anger, bitterness, disappointment; style: ranting, authoritative; average recording, some background noise; genuineness 1.1/6; vocal-burst blend 2.0/10; 18.0s, DE.
DE_1ZMIiMMqy9u_W000002 · in -19.5 dBFS · gain -0.5 dB · emolia-00123
(bitterness, sourness, contempt · brisk, highly aroused, tense, authoritative) Dieser Satz in unserem Grundgesetz ist sozusagen der moralische Imperativ, den die Väter und Mütter unseres Grundgesetzes in dieses Grundgesetz geschrieben haben, weil sie die Lehren aus Nationalsozialismus und der Entmenschlichung der Nazidiktatur gezogen haben.
full caption & clip details
An adult masculine voice; delivery is highly aroused, brisk, tense, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; very clear, almost no disfluency, wide pitch range, normal breath; affect is elated, dominant, fairly guarded; reads as bitterness, sourness, contempt; style: authoritative, dramatic; average recording, some background noise; genuineness 0.9/6; vocal-burst blend 1.5/10; 16.5s, DE.
DE_1ZMIiMMqy9u_W000006 · in -20.6 dBFS · gain +0.6 dB · emolia-00123
Sadness rising ↑identity +0.44 emotion 97 %   sad-Sadness-S3-k3 · #13

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.45, 1.05. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.05.

On the corpus-wide percentile scale those become 0.43, 0.92, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 32 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.334 before conversion and 0.776 after — it rose by 0.442. Neighbour-to-neighbour the worst pair went 0.334 → 0.776. (The earlier render, with segment 1 left raw, scores 0.616 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.549 in the original and +0.531 after conversion — 97 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.80 → 3.08 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0 → 0.4456 → 1.0479normalised 0.451 → 0.940 → 0.983identity cos to seg 1 0.334 → 0.776 +0.442identity cos neighbours 0.334 → 0.776d_b rescored +0.549 → +0.531d_a rescored +0.549 → +0.531d_a mined 1.048d_b mined 0.532min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1c4KtO2dsKEtotal 31.6schain gain +3.0 dBseam step 0.8 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, light breath
(teasing, amusement, intoxication altered states of consciousness · neutral tension, moderately variable, some disfluency, conversational) Es, es, es sind auch, (ahem) äh, andere Sachen gemeint. Das war jetzt bloß ein Beispiel. Aber, (childlike giggle) äh, also ich glaube, sexy Klamotten, meinen Klamottenstil nicht nennen. Einfach nur, oh, ist bequem.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as teasing, amusement, intoxication altered states of consciousness; style: conversational, playful; good recording, quiet background; genuineness 4.5/6; vocal-burst blend 0.9/10; 13.5s, DE.
DE_1c4KtO2dsKE_W000000 · in -12.4 dBFS · gain -7.5 dB · emolia-00167
(fatigue exhaustion, disappointment, confusion · slightly relaxed, fairly steady, some disfluency, monologue) aber abgesehen davon, dass ich generell keine oder ganz selten Poster aufgehängt hab, die irgendwelche Personen gezeigt haben und wenn jetzt nicht, weil die Personen irgendwie
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as fatigue exhaustion, disappointment, confusion; style: monologue, casual; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 0.0/10; 9.8s, DE.
DE_1c4KtO2dsKE_W000005 · in -15.0 dBFS · gain -5.0 dB · emolia-00167
(fatigue exhaustion, longing, sadness · relaxed, fairly steady, frequent disfluency, casual) Ja, natürlich, also, sie da behalten ist dann auch, ich mein, ja, das ist halt auch nicht meine Entscheidung, wer dann noch.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, submissive, neutral openness; reads as fatigue exhaustion, longing, sadness; style: casual, monologue; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 0.5/10; 8.8s, DE.
DE_1c4KtO2dsKE_W000013 · in -14.4 dBFS · gain -5.6 dB · emolia-00167
Sadness rising ↑identity +0.65 emotion 100 %   sad-Sadness-S3-k3 · #14

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.44, 1.58. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.58.

On the corpus-wide percentile scale those become 0.43, 0.92, 1.00 — a total move of +0.57.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 29 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.003 before conversion and 0.654 after — it rose by 0.651. Neighbour-to-neighbour the worst pair went 0.062 → 0.692. (The earlier render, with segment 1 left raw, scores 0.596 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.566 in the original and +0.566 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.94 → 3.06 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0 → 0.4426 → 1.5811normalised 0.451 → 0.939 → 0.997identity cos to seg 1 0.003 → 0.654 +0.651identity cos neighbours 0.062 → 0.692d_b rescored +0.566 → +0.566d_a rescored +0.566 → +0.566d_a mined 1.581d_b mined 0.546min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1eb_VRL8MFItotal 28.8schain gain +2.7 dBseam step 1.0 dBcrossfades 100/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, light breath
(helplessness, longing · measured, subdued, slightly relaxed, formal) habe ich eine Beziehung zu Jesus Christus gefunden.
full caption & clip details
An adult feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness, longing; style: formal, monologue; very good recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.1/10; 3.3s, DE.
DE_1eb_VRL8MFI_W000020 · in -20.4 dBFS · gain +0.4 dB · emolia-00216
(affection, malevolence malice, sexual lust · normal-paced, energised, slightly relaxed, didactic) Möchte dich ermutigen an diesen Tagen, die wir zusammen (low mumble) verbringen und (ahem) das Wort Gottes anschauen, dass wir drauf schauen wollen, wie kannst du ein Leben führen, sodass dein Ende
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very clear, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as affection, malevolence malice, sexual lust; style: didactic, dramatic; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 0.0/10; 13.1s, DE.
DE_1eb_VRL8MFI_W000021 · in -19.6 dBFS · gain -0.4 dB · emolia-00216
(helplessness, distress, infatuation · measured, normally alert, neutral tension, storytelling) Wie komme ich meine Bestimmung? Bitte bete, Daniel, dass ich, dass ich meine Bestimmung erlebe und dass ich meine Berufung ergreife und, und ich spüre dieses Herz, dieses Herz, sie wollen ihr Ziel erreichen bis zum Ende.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as helplessness, distress, infatuation; style: storytelling, dramatic; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 2.9/10; 12.7s, DE.
DE_1eb_VRL8MFI_W000024 · in -18.8 dBFS · gain -1.2 dB · emolia-00216
Sadness rising ↑identity +0.00 emotion 97 %   sad-Sadness-S3-k3 · #15

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 0.67, 1.15. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.15.

On the corpus-wide percentile scale those become 0.43, 0.94, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 34 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.712 before conversion and 0.715 after — it rose by 0.003. Neighbour-to-neighbour the worst pair went 0.723 → 0.696. (The earlier render, with segment 1 left raw, scores 0.727 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.555 in the original and +0.540 after conversion — 97 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.79 → 3.01 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores 0.0 → 0.6689 → 1.1494normalised 0.451 → 0.958 → 0.987identity cos to seg 1 0.712 → 0.715 +0.003identity cos neighbours 0.723 → 0.696d_b rescored +0.555 → +0.540d_a rescored +0.555 → +0.540d_a mined 1.149d_b mined 0.536min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1gZLllrev9ctotal 33.8schain gain +3.2 dBseam step 1.2 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, measured, slightly relaxed, light breath
(normally alert, fairly steady, frequent disfluency, didactic) Wie bereits in einem vergangenen Video erwähnt, ist die Frage nach gut oder schlecht schwer zu beantworten. Diese Frage ist aber sehr wichtig.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.0/10; 8.4s, DE.
DE_1gZLllrev9c_W000000 · in -19.9 dBFS · gain -0.1 dB · emolia-00259
(helplessness, distress, fear · normally alert, fairly steady, some disfluency, monologue) Es kann lediglich in einem sehr starken Konflikt mit meiner eigenen subjektiven Moral stehen.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness, distress, fear; style: monologue, storytelling; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 0.0/10; 6.1s, DE.
DE_1gZLllrev9c_W000005 · in -18.8 dBFS · gain -1.2 dB · emolia-00259
(bitterness, contemplation, sadness · very low-energy, steady, frequent disfluency, monologue) Menschen töten ist schlecht, da es die Menschheit näher an die Nicht-Existenz bringt. Die Besiedlung des weiteren Sonnensystems ist gut, da es die Existenzwahrscheinlichkeit der Menschheit nach Katastrophen von planetaren Ausmaßen erhöht. Selbstmord ist schlecht, da es die Menschheit näher an die Nicht-Existenz bringt.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as bitterness, contemplation, sadness; style: monologue, narration; average recording, quiet background; explicit content; genuineness 2.4/6; vocal-burst blend 0.0/10; 19.7s, DE.
DE_1gZLllrev9c_W000012 · in -21.6 dBFS · gain +1.6 dB · emolia-00259
Sadness rising ↑identity +0.03 emotion 97 %   sad-Sadness-S3-k3 · #16

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.44, 1.11. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.11.

On the corpus-wide percentile scale those become 0.43, 0.92, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 61 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.834 before conversion and 0.865 after — it rose by 0.031. Neighbour-to-neighbour the worst pair went 0.862 → 0.852. (The earlier render, with segment 1 left raw, scores 0.662 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.553 in the original and +0.535 after conversion — 97 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.94 → 3.26 (+0.32) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0 → 0.4373 → 1.1104normalised 0.451 → 0.939 → 0.986identity cos to seg 1 0.834 → 0.865 +0.031identity cos neighbours 0.862 → 0.852d_b rescored +0.553 → +0.535d_a rescored +0.553 → +0.535d_a mined 1.110d_b mined 0.535min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1m-geKYo8XAtotal 60.2schain gain +3.8 dBseam step 1.6 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, balanced body, average recording, quiet background, slow, relaxed, frequent disfluency, somewhat unclear
(anger, fatigue exhaustion, jealousy and envy · subdued, steady, whispered, monologue) Nach viel Nachdenken habe ich mich entschlossen, dieses Video vorzubereiten und es als eine Art Podcast vorzutragen. Das wird euch wehtun und es ist mir nicht egal. Nur Fakten zu meinem Kanal und zu mir.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, slow, relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, fairly guarded; reads as anger, fatigue exhaustion, jealousy and envy; style: whispered, monologue; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 0.7/10; 20.2s, DE.
DE_1m-geKYo8XA_W000000 · in -24.1 dBFS · gain +4.2 dB · emolia-00097
(contempt, anger, bitterness · very low-energy, steady, whispered, monologue) Diese Kommentare sind dann wertvoll für alle. Es sind nachvollziehbare Neuigkeiten, die der Wahrheit dienen. Zu den Fakten. Was tue ich? Grundsätzlich vermittle ich Wissen.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is neutral-toned, slightly dark, rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, fairly guarded; reads as contempt, anger, bitterness; style: whispered, monologue; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.1/10; 15.8s, DE.
DE_1m-geKYo8XA_W000003 · in -23.8 dBFS · gain +3.8 dB · emolia-00097
(disappointment, doubt, sadness · very low-energy, fairly steady, monologue, whispered) Wer nichts weiß, muss glauben. Aus dem Unwissen oder Wissen heraus handelt man entsprechend. Ich hoffte, dass ihr mit dem Wissen, es ist alles gesagt und gezeigt, handelt, und das war falsch. Mein Fehler war, dass ich glaubte, dieser
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, fairly guarded; reads as disappointment, doubt, sadness; style: monologue, whispered; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 1.1/10; 24.7s, DE.
DE_1m-geKYo8XA_W000006 · in -23.3 dBFS · gain +3.3 dB · emolia-00097
Sadness rising ↑identity +0.64 emotion 97 %   sad-Sadness-S3-k3 · #17

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.57, 1.15. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.15.

On the corpus-wide percentile scale those become 0.43, 0.93, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 39 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.027 before conversion and 0.670 after — it rose by 0.643. Neighbour-to-neighbour the worst pair went 0.027 → 0.670. (The earlier render, with segment 1 left raw, scores 0.531 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.554 in the original and +0.536 after conversion — 97 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.81 → 3.03 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0 → 0.5742 → 1.1465normalised 0.451 → 0.950 → 0.987identity cos to seg 1 0.027 → 0.670 +0.643identity cos neighbours 0.027 → 0.670d_b rescored +0.554 → +0.536d_a rescored +0.554 → +0.536d_a mined 1.147d_b mined 0.536min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_28I6DLcapyctotal 38.5schain gain +2.0 dBseam step 2.0 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, slightly relaxed, fairly steady, moderate pitch range
(emotional numbness, contemplation, fear · measured, normally alert, little disfluency, didactic) als sich in einem Betrieb ausbilden zu lassen. Das führt (ahem) zu der Schieflage, die wir jetzt haben.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, contemplation, fear; style: didactic, formal; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.0/10; 6.2s, DE.
DE_28I6DLcapyc_W000000 · in -16.8 dBFS · gain -3.2 dB · emolia-00097
(disappointment, malevolence malice, sadness · normal-paced, normally alert, no disfluency, formal) In dieser Folge erzählt Friedrich Hubert Esser, warum die Berufsausbildung in den vergangenen Jahrzehnten an Ansehen verloren hat und was das mit unserer Gesellschaft macht. Er räumt mit Klischees rund um die Berufsbildung auf und hat auch einen guten Rat für Gymnasium parat.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, malevolence malice, sadness; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 17.3s, DE.
DE_28I6DLcapyc_W000004 · in -16.9 dBFS · gain -3.1 dB · emolia-00097
(jealousy and envy, sadness, concentration · measured, subdued, frequent disfluency, monologue) Und, äh, (ahem) so mit auch in unserem Land mitgestalten und so mitnehmen wir auch vielen, ich sag mal, die Mühen weg, die man aufnehmen muss, wenn man Menschen beispielsweise helfen muss, weil sie
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, sadness, concentration; style: monologue, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 0.0/10; 15.5s, DE.
DE_28I6DLcapyc_W000034 · in -14.8 dBFS · gain -5.2 dB · emolia-00097
Sadness rising ↑identity +0.38 emotion REVERSED   sad-Sadness-S3-k3 · #18

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.52, 1.09. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.09.

On the corpus-wide percentile scale those become 0.43, 0.93, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 28 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.309 before conversion and 0.688 after — it rose by 0.380. Neighbour-to-neighbour the worst pair went 0.373 → 0.689. (The earlier render, with segment 1 left raw, scores 0.598 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

The emotional move did not survive. Re-scored end to end, Sadness moved +0.552 in the original and -0.472 after conversion — it changed direction. On this chain the corrected audio is not an improvement.

Quality. Mean predicted overall quality across the segments went 2.93 → 3.18 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0 → 0.5156 → 1.0928normalised 0.451 → 0.945 → 0.985identity cos to seg 1 0.309 → 0.688 +0.380identity cos neighbours 0.373 → 0.689d_b rescored +0.552 → -0.472d_a rescored +0.552 → -0.472d_a mined 1.093d_b mined 0.534min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_2HgcLfDY9l4total 27.2schain gain +0.2 dBseam step 0.3 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady
(bitterness, jealousy and envy, anger · almost no disfluency, clear, monologue, formal) Braucht es im Rekrutierungsprozess von den Lernenden ein neues Instrument? Mein heutiger Gast ist Patrizia Manducca. Sie ist Ausbildungsverantwortliche beim Schweizer Handelsunternehmen und Lehrbetrieb Pestalozi. Und ich bin Marc Purchet von justy.ch und das ist eine neue Folge von Inside Berufsbildung.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as bitterness, jealousy and envy, anger; style: monologue, formal; good recording, quiet background; genuineness 0.2/6; vocal-burst blend 1.0/10; 17.3s, DE.
DE_2HgcLfDY9l4_W000000 · in -18.0 dBFS · gain -2.0 dB · emolia-00196
(sadness, disappointment, helplessness · no disfluency, clear, casual, formal) Und da haben wir gefunden, wir wollen da andere Sachen ausprobieren, um dem näher zu kommen.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sadness, disappointment, helplessness; style: casual, formal; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 1.1/10; 3.9s, DE.
DE_2HgcLfDY9l4_W000018 · in -19.0 dBFS · gain -1.0 dB · emolia-00196
(disappointment, sadness, longing · some disfluency, average clarity, formal) Am Anfang, natürlich, kommen mir sehr viele Bewerbungen rüber. Wir lassen die Bewerbungen über Justi laufen, über euch.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, sadness, longing; style: formal; average recording, no background noise; genuineness 3.1/6; vocal-burst blend 3.7/10; 6.4s, DE.
DE_2HgcLfDY9l4_W000048 · in -16.4 dBFS · gain -3.6 dB · emolia-00196
Sadness rising ↑identity +0.03 emotion 11 %   sad-Sadness-S3-k3 · #19

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.56, 1.10. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.10.

On the corpus-wide percentile scale those become 0.43, 0.93, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 38 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.830 before conversion and 0.857 after — it rose by 0.027. Neighbour-to-neighbour the worst pair went 0.830 → 0.847. (The earlier render, with segment 1 left raw, scores 0.696 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.552 in the original and +0.062 after conversion — 11 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.61 → 3.09 (+0.48) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0 → 0.5635 → 1.1035normalised 0.451 → 0.949 → 0.986identity cos to seg 1 0.830 → 0.857 +0.027identity cos neighbours 0.830 → 0.847d_b rescored +0.552 → +0.062d_a rescored +0.552 → +0.062d_a mined 1.103d_b mined 0.534min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_2R_mAsRw2RItotal 37.0schain gain +0.8 dBseam step 2.2 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normal-paced, normally alert, some disfluency
(jealousy and envy, contentment, affection · neutral tension, moderately variable, casual, conversational) Und mein erster Weg führt mich dann natürlich, wie bei jedem anderen Menschen denke ich, auch einfach ins Badezimmer. Und so sieht das Ganze aus, wenn ich am Abend zuvor duschen war und meine Haare habe lufttrocknen lassen, dann ist das immer eine ganz lustige Geschichte. Das erste, was ich dann mache, ist natürlich mal kurz auf die Toilette gehen.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as jealousy and envy, contentment, affection; style: casual, conversational; good recording, quiet background; genuineness 4.1/6; vocal-burst blend 3.7/10; 17.3s, DE.
DE_2R_mAsRw2RI_W000006 · in -15.6 dBFS · gain -4.4 dB · emolia-00258
(fatigue exhaustion, pain, disgust · slightly relaxed, fairly steady, casual, monologue) Und auch um meinen Kreislauf in Schwung zu bringen, wenn ich mal keine Zitrone trinke, dann ist es auch mal Apfelessig für die Haut, zur Entgiftung, intervallmäßig und so weiter.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as fatigue exhaustion, pain, disgust; style: casual, monologue; good recording, no background noise; genuineness 2.9/6; vocal-burst blend 0.0/10; 9.4s, DE.
DE_2R_mAsRw2RI_W000021 · in -15.2 dBFS · gain -4.8 dB · emolia-00258
(fatigue exhaustion, longing, contentment · slightly relaxed, fairly steady, casual, monologue) Dehne ich mich am Morgen immer ganz gerne und entspanne nochmal so ein bisschen. Dehne auch meinen Schultern und Nackenbereich dadurch, dass ich ja einen ganzen Tag am Schreibtisch am Computer sitze auf der Arbeit.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as fatigue exhaustion, longing, contentment; style: casual, monologue; good recording, no background noise; genuineness 3.5/6; vocal-burst blend 0.7/10; 10.7s, DE.
DE_2R_mAsRw2RI_W000033 · in -13.6 dBFS · gain -6.4 dB · emolia-00258
Sadness rising ↑identity −0.07 emotion 99 %   sad-Sadness-S3-k3 · #20

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 0.79, 1.39. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.39.

On the corpus-wide percentile scale those become 0.43, 0.96, 0.99 — a total move of +0.56.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 31 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.764 before conversion and 0.696 after — it fell by 0.068. Neighbour-to-neighbour the worst pair went 0.764 → 0.696. (The earlier render, with segment 1 left raw, scores 0.749 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.563 in the original and +0.558 after conversion — 99 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.96 → 3.14 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0 → 0.7896 → 1.3867normalised 0.451 → 0.968 → 0.994identity cos to seg 1 0.764 → 0.696 -0.068identity cos neighbours 0.764 → 0.696d_b rescored +0.563 → +0.558d_a rescored +0.563 → +0.558d_a mined 1.387d_b mined 0.543min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_2YnrA50c9Dototal 30.4schain gain +1.2 dBseam step 2.0 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, measured, normally alert, slightly relaxed
(fear, helplessness, affection · fairly steady, didactic, formal) Angst will dich warnen, will dir helfen, das Überleben zu sichern.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as fear, helplessness, affection; style: didactic, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 5.1s, DE.
DE_2YnrA50c9Do_W000006 · in -19.4 dBFS · gain -0.6 dB · emolia-00060
(contemplation, fear, helplessness · steady, didactic, monologue) Die Vernunft, die kriegt von all dem gar nichts mit, weil unsere Energie ist, währenddem wir in der Angst sind, nicht gleichzeitig in der Vernunft. Wir können nicht vernünftig und ängstlich zugleich sein, entweder oder.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, fear, helplessness; style: didactic, monologue; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.2/10; 14.0s, DE.
DE_2YnrA50c9Do_W000008 · in -21.8 dBFS · gain +1.8 dB · emolia-00060
(sadness, helplessness, fear · fairly steady, didactic, monologue) Je mehr Angst wir haben, desto unvernünftiger werden wir. Und je mehr Vernunft wir haben, desto weniger Angst haben wir. Wenn also die Information über die Vernunft verläuft,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sadness, helplessness, fear; style: didactic, monologue; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.0/10; 11.7s, DE.
DE_2YnrA50c9Do_W000009 · in -19.5 dBFS · gain -0.5 dB · emolia-00060