rescue rule sad-Sadness-S3-k2 — voice-corrected

Sadness under rescue rule S3, k=2. the literal 'not sad -> clearly sad' pair. 90.0 % of clips score at or below zero on this emotion and the largest gap on its normalised axis is 0.449 (WIDER than the 0.25 step cap). This rule found 7,826 chains over 40,000 tracks; the strict rule found 0 at k=3.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_sad-Sadness-S3-k2.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
These are not strict-rule trajectories. They come from a deliberately looser rule, built to recover examples on an emotion the strict rule cannot reach. What rule S3 changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'. What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained. Full explanation →
20chains converted
20segments re-voiced
0.733 → 0.774median worst-to-anchor identity cosine
84 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Sadness rising ↑identity −0.02 emotion 94 %   sad-Sadness-S3-k2 · #1

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 1.09. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.09.

On the corpus-wide percentile scale those become 0.43, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 10 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.716 before conversion and 0.698 after — it fell by 0.017. Neighbour-to-neighbour the worst pair went 0.716 → 0.698. (The earlier render, with segment 1 left raw, scores 0.585 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.551 in the original and +0.520 after conversion — 94 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.83 → 2.98 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.0869normalised 0.451 → 0.985identity cos to seg 1 0.716 → 0.698 -0.017identity cos neighbours 0.716 → 0.698d_b rescored +0.551 → +0.520d_a rescored +0.551 → +0.520d_a mined 1.087d_b mined 0.534min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-WhG-j-UudUtotal 9.6schain gain +2.1 dBseam step 0.2 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, measured, slightly relaxed
(normally alert, fairly steady, little disfluency, authoritative) Hello. Herzlich willkommen aus der Quantum Storm Star Wars Collection. Mein Name ist Deniz und ich möchte euch heute den
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, didactic; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 0.1/10; 6.7s, DE.
DE_-WhG-j-UudU_W000000 · in -19.3 dBFS · gain -0.7 dB · emolia-00238
(helplessness, sadness, longing · subdued, steady, no disfluency, formal) Gummibändern festgemacht, die wir von.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; reads as helplessness, sadness, longing; style: formal, monologue; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 2.0/10; 3.1s, DE.
DE_-WhG-j-UudU_W000024 · in -16.3 dBFS · gain -3.7 dB · emolia-00238
Sadness rising ↑identity +0.16 emotion 95 %   sad-Sadness-S3-k2 · #2

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 1.47. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.47.

On the corpus-wide percentile scale those become 0.44, 0.99 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 33 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.304 before conversion and 0.459 after — it rose by 0.155. Neighbour-to-neighbour the worst pair went 0.304 → 0.459. (The earlier render, with segment 1 left raw, scores 0.369 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.556 in the original and +0.529 after conversion — 95 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.68 → 3.14 (+0.46) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.1855normalised 0.451 → 0.989identity cos to seg 1 0.304 → 0.459 +0.155identity cos neighbours 0.304 → 0.459d_b rescored +0.556 → +0.529d_a rescored +0.556 → +0.529d_a mined 1.185d_b mined 0.538min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-xg8lm69_K8total 33.1schain gain +1.9 dBseam step 1.9 dBcrossfades 100 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-bright, fairly smooth, balanced body, average recording, energised, light breath
(normal-paced, slightly relaxed, fairly steady, authoritative) Und ich rufe auf den Tagesordnungspunkt 11.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very clear, frequent disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, formal; average recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.3/10; 3.2s, DE.
DE_-xg8lm69_K8_W000000 · in -18.2 dBFS · gain -1.8 dB · emolia-00224
(disappointment, distress, impatience and irritability · brisk, neutral tension, moderately variable, monologue) Und da würde ich sagen, liegt es nicht nur daran, dass das Wahlprozedere jetzt irgendwie kompliziert wäre. Ich meine, die Studierenden bekommen die Unterlagen geschickt. Es gibt tagelang Zeit, seine Stimme abzugeben. Ich denke, wir müssen auch überlegen, ob es was damit zu tun hat, (ahem) dass die Entscheidungskompetenzen in der Verfasst Studierendenschaft nicht gerade sehr ausgeprägt sind, um es vorsichtig zu sagen. Also, den ganzen Autonomieprozess, den es gab, sind ja Kompetenzen vom Ministerium an die Hochschulen (ahem) (ahem) verlagert worden, aber eben dort vor allem an die Präsidien und an die Hochschulräte. Und ich finde, man muss vielleicht
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as disappointment, distress, impatience and irritability; style: monologue, dramatic; average recording, some background noise; genuineness 2.1/6; vocal-burst blend 2.7/10; 30.0s, DE.
DE_-xg8lm69_K8_W000052 · in -24.0 dBFS · gain +4.0 dB · emolia-00224
Sadness rising ↑identity +0.03 emotion REVERSED   sad-Sadness-S3-k2 · #3

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 1.24. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.24.

On the corpus-wide percentile scale those become 0.44, 0.99 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 14 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.741 before conversion and 0.775 after — it rose by 0.034. Neighbour-to-neighbour the worst pair went 0.741 → 0.775. (The earlier render, with segment 1 left raw, scores 0.534 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Sadness moved +0.560 in the original and -0.012 after conversion — it changed direction. On this chain the corrected audio is not an improvement.

Quality. Mean predicted overall quality across the segments went 2.45 → 2.87 (+0.42) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.3037normalised 0.451 → 0.992identity cos to seg 1 0.741 → 0.775 +0.034identity cos neighbours 0.741 → 0.775d_b rescored +0.560 → -0.012d_a rescored +0.560 → -0.012d_a mined 1.304d_b mined 0.541min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_02FqQjJyMtytotal 13.3schain gain +3.3 dBseam step 0.2 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an elderly masculine voice · neutral-toned, dark, slightly thin, below-average recording, quiet background, slow, relaxed, steady
(emotional numbness, contemplation · lethargic, narrow pitch range, whispered, monologue) So, ich würd mal sagen, Willkommen zu einer neuen Aufnahme, außerhalb des Streams.
full caption & clip details
An elderly masculine voice; delivery is lethargic, slow, relaxed, steady; timbre is neutral-toned, dark, smooth, slightly thin; slurred, frequent disfluency, narrow pitch range, audible breath; affect is mildly negative, submissive, neutral openness; reads as emotional numbness, contemplation; style: whispered, monologue; below-average recording, quiet background; genuineness 4.3/6; vocal-burst blend 0.0/10; 9.0s, DE.
DE_02FqQjJyMty_W000000 · in -18.4 dBFS · gain -1.6 dB · emolia-00215
(sadness, helplessness, distress · very low-energy, fairly narrow pitch, whispered, casual) In (low mumble) den Videos habe ich ja nicht mal mit dabei. Würde eigentlich Sinn machen.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is neutral-toned, dark, fairly smooth, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, submissive, neutral openness; reads as sadness, helplessness, distress; style: whispered, casual; below-average recording, quiet background; genuineness 5.7/6; vocal-burst blend 0.8/10; 4.6s, DE.
DE_02FqQjJyMty_W000002 · in -17.1 dBFS · gain -2.9 dB · emolia-00215
Sadness rising ↑identity −0.01 emotion 96 %   sad-Sadness-S3-k2 · #4

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 1.02. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.02.

On the corpus-wide percentile scale those become 0.43, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 10 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.695 before conversion and 0.685 after — it fell by 0.010. Neighbour-to-neighbour the worst pair went 0.695 → 0.685. (The earlier render, with segment 1 left raw, scores 0.562 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.548 in the original and +0.526 after conversion — 96 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.98 → 3.03 (+0.05) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.0244normalised 0.451 → 0.982identity cos to seg 1 0.695 → 0.685 -0.010identity cos neighbours 0.695 → 0.685d_b rescored +0.548 → +0.526d_a rescored +0.548 → +0.526d_a mined 1.024d_b mined 0.531min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0FuG3tlx0yQtotal 9.8schain gain +3.9 dBseam step 2.0 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(fairly steady, some disfluency, conversational, authoritative) auch wenn's nicht stattfindet, ein bisschen Lagerstellen, man darf doch hier aufkommen. Lass uns doch jetzt gemeinsam kurz ein Zelt aufstellen. Los geht's!
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: conversational, authoritative; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.0/10; 6.7s, DE.
DE_0FuG3tlx0yQ_W000000 · in -24.3 dBFS · gain +4.3 dB · emolia-00097
(sadness, triumph, helplessness · moderately variable, no disfluency, authoritative, dramatic) Doch selbst noch die besten Jahre sind voll Kummer und Schmerz.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sadness, triumph, helplessness; style: authoritative, dramatic; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.7/10; 3.3s, DE.
DE_0FuG3tlx0yQ_W000006 · in -18.1 dBFS · gain -1.9 dB · emolia-00097
Sadness rising ↑identity +0.11 emotion 83 %   sad-Sadness-S3-k2 · #5

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 1.09. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.09.

On the corpus-wide percentile scale those become 0.43, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 18 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.725 before conversion and 0.835 after — it rose by 0.109. Neighbour-to-neighbour the worst pair went 0.725 → 0.835. (The earlier render, with segment 1 left raw, scores 0.760 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.552 in the original and +0.457 after conversion — 83 % of the delta retained, which is most of it.

Quality. Mean predicted overall quality across the segments went 3.07 → 3.24 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.0928normalised 0.451 → 0.985identity cos to seg 1 0.725 → 0.835 +0.109identity cos neighbours 0.725 → 0.835d_b rescored +0.552 → +0.457d_a rescored +0.552 → +0.457d_a mined 1.093d_b mined 0.534min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0IWgiqWfPNAtotal 17.7schain gain +1.7 dBseam step 1.0 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, some disfluency
(moderately variable, wide pitch range, conversational, playful) Hey Leute, willkommen zurück zu einem neuen Video hier auf Popel mit Zucker. Wie immer fangen wir an mit einem extra Schluck.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: conversational, playful; good recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.4/10; 5.6s, DE.
DE_0IWgiqWfPNA_W000000 · in -17.9 dBFS · gain -2.1 dB · emolia-00195
(fatigue exhaustion, pain, sadness · fairly steady, moderate pitch range, monologue, casual) zu leicht, glaube ich, meiner Meinung nach. Hätte man doppelt überlegen müssen, aber klar, man will die Rache, man will ihn endlich loswerden, nach all der Zeit, und dann laut man halt auch so eine, so eine Täuschung mal eher, aber das ist halt,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, pain, sadness; style: monologue, casual; average recording, no background noise; genuineness 3.1/6; vocal-burst blend 1.1/10; 12.3s, DE.
DE_0IWgiqWfPNA_W000059 · in -20.7 dBFS · gain +0.7 dB · emolia-00195
Sadness rising ↑identity +0.03 emotion 51 %   sad-Sadness-S3-k2 · #6

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.04, 1.00. In the first clip the scorer found no Sadness whatsoever (0.04); by the last it is at 1.00.

On the corpus-wide percentile scale those become 0.87, 0.98 — a total move of +0.10.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 33 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.880 before conversion and 0.909 after — it rose by 0.029. Neighbour-to-neighbour the worst pair went 0.880 → 0.909. (The earlier render, with segment 1 left raw, scores 0.728 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.104 in the original and +0.053 after conversion — 51 % of the delta retained.

Quality. Mean predicted overall quality across the segments went 2.84 → 3.25 (+0.42) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.043 → 1.002normalised 0.910 → 0.981identity cos to seg 1 0.880 → 0.909 +0.029identity cos neighbours 0.880 → 0.909d_b rescored +0.104 → +0.053d_a rescored +0.104 → +0.053d_a mined 0.959d_b mined 0.071min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0KZXkf63R18total 32.7schain gain +2.5 dBseam step 0.2 dBcrossfades 100 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a middle-aged feminine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, very low-energy, fairly steady
(doubt, confusion, pain · normal-paced, slightly relaxed, some disfluency, casual) Ist es überhaupt richtig, das Setting, was mir jetzt gerade angeboten wird? Und genau darum geht es in diesem Video. Und zwar um die Aspekte, die man berücksichtigen kann oder sollte, bevor man sich dafür oder dagegen entscheidet.
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as doubt, confusion, pain; style: casual, monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.0/10; 16.3s, DE.
DE_0KZXkf63R18_W000001 · in -23.9 dBFS · gain +3.9 dB · emolia-00254
(sadness, helplessness, contemplation · slow, relaxed, frequent disfluency, monologue) Leider ist zurzeit (low mumble) auch unter Therapeuten und Trauma-Behandlenden die Situation folgend, dass (low mumble) bei vielen nicht ausreichend, ja, Wissen und Kenntnisse da sind.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, submissive, neutral openness; reads as sadness, helplessness, contemplation; style: monologue, ASMR; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 0.0/10; 16.5s, DE.
DE_0KZXkf63R18_W000004 · in -20.6 dBFS · gain +0.6 dB · emolia-00254
Sadness rising ↑identity −0.07 emotion 16 %   sad-Sadness-S3-k2 · #7

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 1.12. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.12.

On the corpus-wide percentile scale those become 0.44, 0.98 — a total move of +0.54.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 11 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.802 before conversion and 0.729 after — it fell by 0.073. Neighbour-to-neighbour the worst pair went 0.802 → 0.729. (The earlier render, with segment 1 left raw, scores 0.493 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.556 in the original and +0.088 after conversion — 16 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.46 → 2.81 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.1709normalised 0.451 → 0.988identity cos to seg 1 0.802 → 0.729 -0.073identity cos neighbours 0.802 → 0.729d_b rescored +0.556 → +0.088d_a rescored +0.556 → +0.088d_a mined 1.171d_b mined 0.537min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0PbfumbfjeQtotal 10.6schain gain +2.3 dBseam step 0.0 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, slightly relaxed, some disfluency, average clarity, moderate pitch range, light breath
(fatigue exhaustion · normal-paced, normally alert, moderately variable, casual) Karten und so weiter. Ich mach's gleich auf und zeig's euch richtig, aber ich zeig's einfach nur so kurz und dann.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion; style: casual, conversational; below-average recording, quiet background; genuineness 5.0/6; vocal-burst blend 2.8/10; 5.3s, DE.
DE_0PbfumbfjeQ_W000000 · in -19.7 dBFS · gain -0.3 dB · emolia-00161
(helplessness, sadness, distress · measured, subdued, fairly steady, monologue) Jetzt werde ich auf jeden Fall keine weiteren mehr bestellen. Das muss man jetzt erstmal alles verarbeiten, was ich hier.
full caption & clip details
An adult feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as helplessness, sadness, distress; style: monologue, ASMR; good recording, no background noise; genuineness 3.1/6; vocal-burst blend 0.7/10; 5.5s, DE.
DE_0PbfumbfjeQ_W000005 · in -18.4 dBFS · gain -1.6 dB · emolia-00161
Sadness rising ↑identity +0.00 emotion 100 %   sad-Sadness-S3-k2 · #8

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 1.01. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.01.

On the corpus-wide percentile scale those become 0.43, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 15 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.913 before conversion and 0.915 after — it rose by 0.002. Neighbour-to-neighbour the worst pair went 0.913 → 0.915. (The earlier render, with segment 1 left raw, scores 0.888 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.547 in the original and +0.549 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.61 → 2.96 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.0098normalised 0.451 → 0.982identity cos to seg 1 0.913 → 0.915 +0.002identity cos neighbours 0.913 → 0.915d_b rescored +0.547 → +0.549d_a rescored +0.547 → +0.549d_a mined 1.010d_b mined 0.530min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0S219bZ30GQtotal 14.8schain gain +3.1 dBseam step 1.6 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(steady, newsreading, formal) Zusammen wollen wir uns verschiedene Begriffe aus dem queeren und feministischen Spektrum anschauen, Hintergründe zu den Themen beleuchten und Fragen klären.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.2/10; 7.3s, DE.
DE_0S219bZ30GQ_W000000 · in -22.0 dBFS · gain +2.0 dB · emolia-00097
(contempt, sadness, distress · fairly steady, formal, newsreading) Die Begriffe Zwitter und Hermaphrodit sind für die meisten intersexuellen Menschen sehr beleidigend und verletzend. Du solltest sie also nicht benutzen.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, sadness, distress; style: formal, newsreading; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.2/10; 7.7s, DE.
DE_0S219bZ30GQ_W000012 · in -21.3 dBFS · gain +1.3 dB · emolia-00097
Sadness rising ↑identity −0.00 emotion REVERSED   sad-Sadness-S3-k2 · #9

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 1.30. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.30.

On the corpus-wide percentile scale those become 0.43, 0.99 — a total move of +0.56.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 25 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.805 before conversion and 0.801 after — it fell by 0.004. Neighbour-to-neighbour the worst pair went 0.805 → 0.801. (The earlier render, with segment 1 left raw, scores 0.701 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The emotional move did not survive. Re-scored end to end, Sadness moved +0.560 in the original and -0.044 after conversion — it changed direction. On this chain the corrected audio is not an improvement.

Quality. Mean predicted overall quality across the segments went 2.69 → 3.07 (+0.38) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.2949normalised 0.451 → 0.992identity cos to seg 1 0.805 → 0.801 -0.004identity cos neighbours 0.805 → 0.801d_b rescored +0.560 → -0.044d_a rescored +0.560 → -0.044d_a mined 1.295d_b mined 0.541min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0fB4TCjiFIktotal 24.4schain gain +1.9 dBseam step 0.1 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(concentration, hope enthusiasm optimism, contemplation · monologue, formal) Ihr Lieben, (low mumble) für uns Grüne war schon immer klar, die Zukunft, die wir wollen, die müssen wir auch selber erschaffen. Und auch wenn die Folgen der Corona-Krise uns noch lange beschäftigen werden, wenn nicht jetzt alle Probleme, alle Krisen und alle Herausforderungen gleichzeitig angeht, der verspielt die Zukunft der zukünftigen Generation.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, hope enthusiasm optimism, contemplation; style: monologue, formal; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.0/10; 18.3s, DE.
DE_0fB4TCjiFIk_W000000 · in -15.9 dBFS · gain -4.1 dB · emolia-00077
(pride, helplessness, sadness · didactic, monologue) Und wir sind keine kleine Partei mehr. Wir sind wahnsinnig gewachsen. Als Nina und ich als Landesvorsitzende gewählt wurden.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, helplessness, sadness; style: didactic, monologue; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.0/10; 6.3s, DE.
DE_0fB4TCjiFIk_W000010 · in -14.9 dBFS · gain -5.1 dB · emolia-00077
Sadness rising ↑identity −0.01 emotion REVERSED   sad-Sadness-S3-k2 · #10

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 1.12. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.12.

On the corpus-wide percentile scale those become 0.43, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 11 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.780 before conversion and 0.773 after — it fell by 0.007. Neighbour-to-neighbour the worst pair went 0.780 → 0.773. (The earlier render, with segment 1 left raw, scores 0.596 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.553 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.56 → 2.82 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.1172normalised 0.451 → 0.986identity cos to seg 1 0.780 → 0.773 -0.007identity cos neighbours 0.780 → 0.773d_b rescored +0.553 → +0.000d_a rescored +0.553 → +0.000d_a mined 1.117d_b mined 0.535min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0ytNrKylOrctotal 10.5schain gain +2.0 dBseam step 3.7 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an elderly somewhat feminine voice · slightly warm, balanced body, good recording, no background noise, slow, very low-energy, relaxed, steady
(sexual lust, fatigue exhaustion, infatuation · no disfluency, slurred, narrow pitch range, whispered) denn die Energien öffnen deinen Geist.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, smooth, balanced body; slurred, no disfluency, narrow pitch range, audible breath; affect is mildly positive, submissive, neutral openness; reads as sexual lust, fatigue exhaustion, infatuation; style: whispered, monologue; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 5.2/10; 3.9s, DE.
DE_0ytNrKylOrc_W000002 · in -13.5 dBFS · gain -6.5 dB · emolia-00215
(distress, fear, sadness · frequent disfluency, somewhat unclear, fairly narrow pitch, ASMR) dass du als Mensch hier auf der Erde das vergessen hast.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, submissive, vulnerable; reads as distress, fear, sadness; style: ASMR, monologue; good recording, no background noise; genuineness 2.8/6; vocal-burst blend 4.2/10; 6.8s, DE.
DE_0ytNrKylOrc_W000003 · in -15.8 dBFS · gain -4.2 dB · emolia-00215
Sadness rising ↑identity −0.03 emotion 100 %   sad-Sadness-S3-k2 · #11

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 1.67. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.67.

On the corpus-wide percentile scale those become 0.43, 1.00 — a total move of +0.57.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 11 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.884 before conversion and 0.849 after — it fell by 0.035. Neighbour-to-neighbour the worst pair went 0.884 → 0.849. (The earlier render, with segment 1 left raw, scores 0.792 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.567 in the original and +0.565 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.79 → 2.98 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.668normalised 0.451 → 0.998identity cos to seg 1 0.884 → 0.849 -0.035identity cos neighbours 0.884 → 0.849d_b rescored +0.567 → +0.565d_a rescored +0.567 → +0.565d_a mined 1.668d_b mined 0.546min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_12MSkDIQFvktotal 10.7schain gain -2.2 dBseam step 1.7 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged somewhat feminine voice · neutral-toned, slightly dark, fairly smooth, balanced body, average recording, quiet background, very low-energy, relaxed
(fatigue exhaustion, confusion, doubt · measured, average clarity, fairly narrow pitch, ASMR) Nachbarin auch mitnehmen und weiter, weiter viertn.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, submissive, neutral openness; reads as fatigue exhaustion, confusion, doubt; style: ASMR, whispered; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 0.2/10; 4.5s, DE.
DE_12MSkDIQFvk_W000006 · in -16.6 dBFS · gain -3.4 dB · emolia-00003
(sadness, pain, helplessness · slow, somewhat unclear, moderate pitch range, ASMR) Kleine Kinder alle angeblie, alleine geblieben waren. Der älteste war, war sechs Jahre.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, submissive, neutral openness; reads as sadness, pain, helplessness; style: ASMR, whispered; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 0.8/10; 6.3s, DE.
DE_12MSkDIQFvk_W000009 · in -17.0 dBFS · gain -3.0 dB · emolia-00003
Sadness rising ↑identity −0.05 emotion 84 %   sad-Sadness-S3-k2 · #12

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 1.02. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.02.

On the corpus-wide percentile scale those become 0.44, 0.98 — a total move of +0.54.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 12 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.907 before conversion and 0.860 after — it fell by 0.048. Neighbour-to-neighbour the worst pair went 0.907 → 0.860. (The earlier render, with segment 1 left raw, scores 0.840 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.549 in the original and +0.464 after conversion — 84 % of the delta retained, which is most of it.

Quality. Mean predicted overall quality across the segments went 3.03 → 3.18 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.0508normalised 0.451 → 0.983identity cos to seg 1 0.907 → 0.860 -0.048identity cos neighbours 0.907 → 0.860d_b rescored +0.549 → +0.464d_a rescored +0.549 → +0.464d_a mined 1.051d_b mined 0.532min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_163DYBR_hRgtotal 11.8schain gain +2.1 dBseam step 0.8 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(steady, formal, monologue) Startschuss für DSDS. Nicht mehr lange, und Florian Silbereisen-Fans kommen auch bei RTL auf ihre Kosten
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.9s, DE.
DE_163DYBR_hRg_W000000 · in -16.4 dBFS · gain -3.6 dB · emolia-00161
(confusion, helplessness, sadness · fairly steady, storytelling, formal) Ich sitze hier und weiß genauso wenig wie ihr, was heute passiert.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as confusion, helplessness, sadness; style: storytelling, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.5/10; 4.1s, DE.
DE_163DYBR_hRg_W000011 · in -16.4 dBFS · gain -3.6 dB · emolia-00161
Sadness rising ↑identity −0.02 emotion REVERSED   sad-Sadness-S3-k2 · #13

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.01, 1.21. In the first clip the scorer found no Sadness whatsoever (0.01); by the last it is at 1.21.

On the corpus-wide percentile scale those become 0.87, 0.99 — a total move of +0.12.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 15 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.479 before conversion and 0.456 after — it fell by 0.023. Neighbour-to-neighbour the worst pair went 0.479 → 0.456. (The earlier render, with segment 1 left raw, scores 0.517 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

The emotional move did not survive. Re-scored end to end, Sadness moved +0.122 in the original and -0.531 after conversion — it changed direction. On this chain the corrected audio is not an improvement.

Quality. Mean predicted overall quality across the segments went 2.63 → 2.86 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0143 → 1.209normalised 0.905 → 0.989identity cos to seg 1 0.479 → 0.456 -0.023identity cos neighbours 0.479 → 0.456d_b rescored +0.122 → -0.531d_a rescored +0.122 → -0.531d_a mined 1.195d_b mined 0.084min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_17Ls9FmMIuktotal 14.3schain gain +3.3 dBseam step 0.8 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged feminine voice · neutral-bright, balanced body, average recording
(shame · measured, normally alert, slightly relaxed, storytelling) sollst fürs Vaterland stechen und schießen.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame; style: storytelling, casual; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 0.4/10; 4.1s, DE.
DE_17Ls9FmMIuk_W000001 · in -20.7 dBFS · gain +0.7 dB · emolia-00245
(longing, sadness, infatuation · slow, energised, relaxed, cartoonish) Drink, mein Söhnchen, von meiner Brust. Drink, dann wirst du ein starker Held. Ziehst mit den andern hinaus ins Feld.
full caption & clip details
A child feminine voice; delivery is energised, slow, relaxed, moderately variable; timbre is cool, neutral-bright, smooth, balanced body; slurred, almost no disfluency, very wide pitch range, no audible breath; affect is positive, slightly submissive, neutral openness; reads as longing, sadness, infatuation; style: cartoonish, storytelling; average recording, some background noise; genuineness 0.9/6; vocal-burst blend 2.7/10; 10.5s, DE.
DE_17Ls9FmMIuk_W000002 · in -18.2 dBFS · gain -1.8 dB · emolia-00245
Sadness rising ↑identity +0.01 emotion 94 %   sad-Sadness-S3-k2 · #14

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 1.38. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.38.

On the corpus-wide percentile scale those become 0.43, 0.99 — a total move of +0.56.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 34 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.812 before conversion and 0.826 after — it rose by 0.014. Neighbour-to-neighbour the worst pair went 0.812 → 0.826. (The earlier render, with segment 1 left raw, scores 0.728 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.562 in the original and +0.527 after conversion — 94 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.92 → 3.11 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.3838normalised 0.451 → 0.994identity cos to seg 1 0.812 → 0.826 +0.014identity cos neighbours 0.812 → 0.826d_b rescored +0.562 → +0.527d_a rescored +0.562 → +0.527d_a mined 1.384d_b mined 0.543min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1A7Boronwiktotal 33.6schain gain +2.5 dBseam step 0.6 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an elderly somewhat feminine voice · slightly warm, neutral-bright, smooth, balanced body, no background noise, slow, very low-energy, relaxed
(contentment, affection, relief · no disfluency, ASMR, monologue) Es ist mir ein Bedürfnis, heute mit dir zusammen, mit euch allen zusammen.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, neutral-bright, smooth, balanced body; clear, no disfluency, narrow pitch range, audible breath; affect is mildly positive, submissive, neutral openness; reads as contentment, affection, relief; style: ASMR, monologue; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.4/10; 8.7s, DE.
DE_1A7Boronwik_W000000 · in -18.1 dBFS · gain -1.9 dB · emolia-00254
(contemplation, sadness, awe · frequent disfluency, whispered, ASMR) Ein Gebet, welches das transformiert, was so schwer auf unseren Schultern liegt. Ein Gebet, was sein Licht über die ganze Erde erstrahlen lässt und die Menschen dazu bringt, den Frieden wieder einzuladen.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, neutral-bright, smooth, balanced body; clear, frequent disfluency, narrow pitch range, audible breath; affect is mildly positive, submissive, vulnerable; reads as contemplation, sadness, awe; style: whispered, ASMR; very good recording, no background noise; genuineness 1.3/6; vocal-burst blend 1.0/10; 25.1s, DE.
DE_1A7Boronwik_W000001 · in -18.8 dBFS · gain -1.2 dB · emolia-00254
Sadness rising ↑identity +0.29 emotion 9 %   sad-Sadness-S3-k2 · #15

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 1.20. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.20.

On the corpus-wide percentile scale those become 0.43, 0.99 — a total move of +0.56.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 10 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.086 before conversion and 0.374 after — it rose by 0.288. Neighbour-to-neighbour the worst pair went 0.086 → 0.374. (The earlier render, with segment 1 left raw, scores 0.419 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.557 in the original and +0.048 after conversion — 9 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.56 → 2.77 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.1973normalised 0.451 → 0.989identity cos to seg 1 0.086 → 0.374 +0.288identity cos neighbours 0.086 → 0.374d_b rescored +0.557 → +0.048d_a rescored +0.557 → +0.048d_a mined 1.197d_b mined 0.538min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1Pd71w8hAhItotal 9.4schain gain +1.2 dBseam step 0.6 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a child feminine voice · neutral-toned, neutral-bright, fairly smooth, average recording, quiet background, moderately variable, wide pitch range
(impatience and irritability, anger · normal-paced, normally alert, slightly relaxed, conversational) Ach so. Auf der Suche nach dem a priori.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as impatience and irritability, anger; style: conversational, casual; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 1.6/10; 3.1s, DE.
DE_1Pd71w8hAhI_W000001 · in -15.7 dBFS · gain -4.3 dB · emolia-00077
(helplessness, fatigue exhaustion, distress · measured, very low-energy, neutral tension, conversational) (low mumble) euhm, also, ich hab so das Gefühl, ich lass mir da auch Zeit, ja, das drängt mich niemand. (breathy giggle)
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, vulnerable; reads as helplessness, fatigue exhaustion, distress; style: conversational, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 1.9/10; 6.5s, DE.
DE_1Pd71w8hAhI_W000014 · in -17.7 dBFS · gain -2.3 dB · emolia-00077
Sadness rising ↑identity +0.17 emotion 96 %   sad-Sadness-S3-k2 · #16

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 1.15. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.15.

On the corpus-wide percentile scale those become 0.43, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 20 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.316 before conversion and 0.489 after — it rose by 0.173. Neighbour-to-neighbour the worst pair went 0.316 → 0.489. (The earlier render, with segment 1 left raw, scores 0.452 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.555 in the original and +0.535 after conversion — 96 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.86 → 3.05 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.1514normalised 0.451 → 0.987identity cos to seg 1 0.316 → 0.489 +0.173identity cos neighbours 0.316 → 0.489d_b rescored +0.555 → +0.535d_a rescored +0.555 → +0.535d_a mined 1.151d_b mined 0.536min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1ZMIiMMqy9utotal 19.4schain gain +2.7 dBseam step 2.6 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-bright, fairly smooth, balanced body, average recording
(normal-paced, normally alert, slightly relaxed, authoritative) Antrag der Fraktion Bündnis 90, die GRÜNEN.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, no audible breath; affect is neutral, dominant, slightly guarded; no dominant emotion; style: authoritative, formal; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.6/10; 3.1s, DE.
DE_1ZMIiMMqy9u_W000000 · in -20.9 dBFS · gain +0.9 dB · emolia-00123
(bitterness, sourness, contempt · brisk, highly aroused, tense, authoritative) Dieser Satz in unserem Grundgesetz ist sozusagen der moralische Imperativ, den die Väter und Mütter unseres Grundgesetzes in dieses Grundgesetz geschrieben haben, weil sie die Lehren aus Nationalsozialismus und der Entmenschlichung der Nazidiktatur gezogen haben.
full caption & clip details
An adult masculine voice; delivery is highly aroused, brisk, tense, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; very clear, almost no disfluency, wide pitch range, normal breath; affect is elated, dominant, fairly guarded; reads as bitterness, sourness, contempt; style: authoritative, dramatic; average recording, some background noise; genuineness 0.9/6; vocal-burst blend 1.5/10; 16.5s, DE.
DE_1ZMIiMMqy9u_W000006 · in -20.6 dBFS · gain +0.6 dB · emolia-00123
Sadness rising ↑identity +0.40 emotion 99 %   sad-Sadness-S3-k2 · #17

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 1.05. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.05.

On the corpus-wide percentile scale those become 0.43, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 22 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.390 before conversion and 0.790 after — it rose by 0.400. Neighbour-to-neighbour the worst pair went 0.390 → 0.790. (The earlier render, with segment 1 left raw, scores 0.655 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.549 in the original and +0.542 after conversion — 99 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.79 → 3.09 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.0479normalised 0.451 → 0.983identity cos to seg 1 0.390 → 0.790 +0.400identity cos neighbours 0.390 → 0.790d_b rescored +0.549 → +0.542d_a rescored +0.549 → +0.542d_a mined 1.048d_b mined 0.532min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1c4KtO2dsKEtotal 22.1schain gain +2.7 dBseam step 0.7 dBcrossfades 100 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, light breath
(teasing, amusement, intoxication altered states of consciousness · neutral tension, moderately variable, some disfluency, conversational) Es, es, es sind auch, (ahem) äh, andere Sachen gemeint. Das war jetzt bloß ein Beispiel. Aber, (childlike giggle) äh, also ich glaube, sexy Klamotten, meinen Klamottenstil nicht nennen. Einfach nur, oh, ist bequem.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as teasing, amusement, intoxication altered states of consciousness; style: conversational, playful; good recording, quiet background; genuineness 4.5/6; vocal-burst blend 0.9/10; 13.5s, DE.
DE_1c4KtO2dsKE_W000000 · in -12.4 dBFS · gain -7.5 dB · emolia-00167
(fatigue exhaustion, longing, sadness · relaxed, fairly steady, frequent disfluency, casual) Ja, natürlich, also, sie da behalten ist dann auch, ich mein, ja, das ist halt auch nicht meine Entscheidung, wer dann noch.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, submissive, neutral openness; reads as fatigue exhaustion, longing, sadness; style: casual, monologue; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 0.5/10; 8.8s, DE.
DE_1c4KtO2dsKE_W000013 · in -14.4 dBFS · gain -5.6 dB · emolia-00167
Sadness rising ↑identity +0.49 emotion 78 %   sad-Sadness-S3-k2 · #18

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 1.85. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.85.

On the corpus-wide percentile scale those become 0.44, 1.00 — a total move of +0.56.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 12 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.267 before conversion and 0.757 after — it rose by 0.490. Neighbour-to-neighbour the worst pair went 0.267 → 0.757. (The earlier render, with segment 1 left raw, scores 0.662 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.568 in the original and +0.444 after conversion — 78 % of the delta retained, which is most of it.

Quality. Mean predicted overall quality across the segments went 2.88 → 3.05 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.751normalised 0.451 → 0.998identity cos to seg 1 0.267 → 0.757 +0.490identity cos neighbours 0.267 → 0.757d_b rescored +0.568 → +0.444d_a rescored +0.568 → +0.444d_a mined 1.751d_b mined 0.547min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1eb_VRL8MFItotal 11.5schain gain +1.0 dBseam step 1.8 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normal-paced, normally alert, fairly steady
(affection · slightly relaxed, little disfluency, clear, authoritative) Tage der Ermutigung. Ich freue mich, dass du wieder eingeschaltet hast, dass du wieder dabei bist.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as affection; style: authoritative, storytelling; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.0/10; 4.9s, DE.
DE_1eb_VRL8MFI_W000000 · in -19.3 dBFS · gain -0.7 dB · emolia-00216
(pain, relief, helplessness · neutral tension, some disfluency, average clarity, conversational) Ich wurde so krank, dass ich nicht mehr laufen konnte, dass ich mich nicht mehr bewegen konnte und ich große Schmerzen hatte im ganzen Körper.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as pain, relief, helplessness; style: conversational, casual; good recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.3/10; 6.9s, DE.
DE_1eb_VRL8MFI_W000004 · in -20.9 dBFS · gain +0.9 dB · emolia-00216
Sadness rising ↑identity +0.02 emotion 99 %   sad-Sadness-S3-k2 · #19

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is 0.00, 1.15. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.15.

On the corpus-wide percentile scale those become 0.43, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 28 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.712 before conversion and 0.734 after — it rose by 0.021. Neighbour-to-neighbour the worst pair went 0.712 → 0.734. (The earlier render, with segment 1 left raw, scores 0.721 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.555 in the original and +0.550 after conversion — 99 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.87 → 3.11 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.1494normalised 0.451 → 0.987identity cos to seg 1 0.712 → 0.734 +0.021identity cos neighbours 0.712 → 0.734d_b rescored +0.555 → +0.550d_a rescored +0.555 → +0.550d_a mined 1.149d_b mined 0.536min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1gZLllrev9ctotal 27.8schain gain +2.8 dBseam step 0.5 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, measured, slightly relaxed, frequent disfluency
(normally alert, fairly steady, didactic, monologue) Wie bereits in einem vergangenen Video erwähnt, ist die Frage nach gut oder schlecht schwer zu beantworten. Diese Frage ist aber sehr wichtig.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.0/10; 8.4s, DE.
DE_1gZLllrev9c_W000000 · in -19.9 dBFS · gain -0.1 dB · emolia-00259
(bitterness, contemplation, sadness · very low-energy, steady, monologue, narration) Menschen töten ist schlecht, da es die Menschheit näher an die Nicht-Existenz bringt. Die Besiedlung des weiteren Sonnensystems ist gut, da es die Existenzwahrscheinlichkeit der Menschheit nach Katastrophen von planetaren Ausmaßen erhöht. Selbstmord ist schlecht, da es die Menschheit näher an die Nicht-Existenz bringt.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as bitterness, contemplation, sadness; style: monologue, narration; average recording, quiet background; explicit content; genuineness 2.4/6; vocal-burst blend 0.0/10; 19.7s, DE.
DE_1gZLllrev9c_W000012 · in -21.6 dBFS · gain +1.6 dB · emolia-00259
Sadness rising ↑identity +0.02 emotion 19 %   sad-Sadness-S3-k2 · #20

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Sadness.

The raw scorer output across the chain is -0.00, 1.11. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 1.11.

On the corpus-wide percentile scale those become 0.43, 0.98 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 45 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.834 before conversion and 0.849 after — it rose by 0.015. Neighbour-to-neighbour the worst pair went 0.834 → 0.849. (The earlier render, with segment 1 left raw, scores 0.655 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.553 in the original and +0.104 after conversion — 19 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.93 → 3.24 (+0.31) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.1104normalised 0.451 → 0.986identity cos to seg 1 0.834 → 0.849 +0.015identity cos neighbours 0.834 → 0.849d_b rescored +0.553 → +0.104d_a rescored +0.553 → +0.104d_a mined 1.110d_b mined 0.535min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1m-geKYo8XAtotal 44.7schain gain +4.0 dBseam step 2.4 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, slightly rough, balanced body, average recording, quiet background, slow, relaxed, frequent disfluency
(anger, fatigue exhaustion, jealousy and envy · subdued, steady, whispered, monologue) Nach viel Nachdenken habe ich mich entschlossen, dieses Video vorzubereiten und es als eine Art Podcast vorzutragen. Das wird euch wehtun und es ist mir nicht egal. Nur Fakten zu meinem Kanal und zu mir.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, slow, relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, fairly guarded; reads as anger, fatigue exhaustion, jealousy and envy; style: whispered, monologue; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 0.7/10; 20.2s, DE.
DE_1m-geKYo8XA_W000000 · in -24.1 dBFS · gain +4.2 dB · emolia-00097
(disappointment, doubt, sadness · very low-energy, fairly steady, monologue, whispered) Wer nichts weiß, muss glauben. Aus dem Unwissen oder Wissen heraus handelt man entsprechend. Ich hoffte, dass ihr mit dem Wissen, es ist alles gesagt und gezeigt, handelt, und das war falsch. Mein Fehler war, dass ich glaubte, dieser
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, fairly guarded; reads as disappointment, doubt, sadness; style: monologue, whispered; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 1.1/10; 24.7s, DE.
DE_1m-geKYo8XA_W000006 · in -23.3 dBFS · gain +3.3 dB · emolia-00097