rescue rule sad-Disappointment-S3-k2 — voice-corrected

Disappointment under rescue rule S3, k=2. generalisation check, gap 0.422. 84.3 % of clips score at or below zero on this emotion and the largest gap on its normalised axis is 0.422 (WIDER than the 0.25 step cap). This rule found 13,495 chains over 40,000 tracks; the strict rule found 0 at k=3.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_sad-Disappointment-S3-k2.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
These are not strict-rule trajectories. They come from a deliberately looser rule, built to recover examples on an emotion the strict rule cannot reach. What rule S3 changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'. What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained. Full explanation →
20chains converted
20segments re-voiced
0.754 → 0.818median worst-to-anchor identity cosine
93 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Disappointment rising ↑identity +0.07 emotion 97 %   sad-Disappointment-S3-k2 · #1

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.00, 1.01. In the first clip the scorer found no Disappointment whatsoever (0.00); by the last it is at 1.01.

On the corpus-wide percentile scale those become 0.39, 0.94 — a total move of +0.54.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 20 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.721 before conversion and 0.794 after — it rose by 0.073. Neighbour-to-neighbour the worst pair went 0.721 → 0.794. (The earlier render, with segment 1 left raw, scores 0.704 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.543 in the original and +0.525 after conversion — 97 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.93 → 3.15 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.0107normalised 0.422 → 0.960identity cos to seg 1 0.721 → 0.794 +0.073identity cos neighbours 0.721 → 0.794d_b rescored +0.543 → +0.525d_a rescored +0.543 → +0.525d_a mined 1.011d_b mined 0.538min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_--z5fTsHDactotal 19.1schain gain +0.4 dBseam step 0.7 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, slightly relaxed, light breath
(normal-paced, normally alert, fairly steady, monologue) Und zwar stellt sich in unserem Fall die Frage, ob denn der Familienrichter, Dr. Markus Bühler,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 0.0/10; 6.2s, DE.
DE_--z5fTsHDac_W000000 · in -18.5 dBFS · gain -1.5 dB · emolia-00037
(fear, concentration, disgust · measured, subdued, steady, didactic) dass dieser Herr Doktor Markus Bühler da irgendwas dran ändern wird, weil das derart starke Muster sind, die er wahrscheinlich schlecht reflektieren kann und auch schlecht verstehen kann.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, concentration, disgust; style: didactic, monologue; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.4/10; 13.2s, DE.
DE_--z5fTsHDac_W000018 · in -18.3 dBFS · gain -1.7 dB · emolia-00037
Disappointment rising ↑identity +0.04 emotion 102 %   sad-Disappointment-S3-k2 · #2

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.00, 1.26. In the first clip the scorer found no Disappointment whatsoever (0.00); by the last it is at 1.26.

On the corpus-wide percentile scale those become 0.39, 0.98 — a total move of +0.59.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 30 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.727 before conversion and 0.770 after — it rose by 0.044. Neighbour-to-neighbour the worst pair went 0.727 → 0.770. (The earlier render, with segment 1 left raw, scores 0.742 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.587 in the original and +0.597 after conversion — 102 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.86 → 3.13 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.2627normalised 0.422 → 0.989identity cos to seg 1 0.727 → 0.770 +0.044identity cos neighbours 0.727 → 0.770d_b rescored +0.587 → +0.597d_a rescored +0.587 → +0.597d_a mined 1.263d_b mined 0.568min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-05ZmfVq_H8total 30.0schain gain +1.2 dBseam step 0.2 dBcrossfades 150 ms
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, normal-paced, normally alert, slightly relaxed
(somewhat unclear, monologue, didactic) Und, (ahem) äh, deshalb unterstützen wir die Bundesregierung bei ihrer klaren Haltung, äh, (ahem) gegenüber Russland.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 0.2/10; 6.1s, DE.
DE_-05ZmfVq_H8_W000001 · in -17.1 dBFS · gain -2.9 dB · emolia-00037
(concentration, disappointment, bitterness · average clarity, monologue, formal) Aber wir halten es mit dieser klaren Haltung nicht für vereinbar, wenn beispielsweise große Infrastrukturprojekte wie die Pipeline Nord Stream 2 vorangetrieben werden, als wäre nichts gewesen. Wir sind nicht für ein prinzipielles Aus- oder einen sofortigen (ahem) Stopp (ahem) dieses Vorhabens, aber es muss ein Moratorium (ahem) geben, bis die Vorgänge
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as concentration, disappointment, bitterness; style: monologue, formal; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 0.0/10; 24.2s, DE.
DE_-05ZmfVq_H8_W000002 · in -16.1 dBFS · gain -3.9 dB · emolia-00037
Disappointment rising ↑identity +0.00 emotion 90 %   sad-Disappointment-S3-k2 · #3

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.00, 1.30. In the first clip the scorer found no Disappointment whatsoever (0.00); by the last it is at 1.30.

On the corpus-wide percentile scale those become 0.39, 0.99 — a total move of +0.59.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 24 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.793 before conversion and 0.797 after — it rose by 0.004. Neighbour-to-neighbour the worst pair went 0.793 → 0.797. (The earlier render, with segment 1 left raw, scores 0.719 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.591 in the original and +0.534 after conversion — 90 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 3.04 → 3.13 (+0.09) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.3027normalised 0.422 → 0.991identity cos to seg 1 0.793 → 0.797 +0.004identity cos neighbours 0.793 → 0.797d_b rescored +0.591 → +0.534d_a rescored +0.591 → +0.534d_a mined 1.303d_b mined 0.570min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-6F1Ss25DJQtotal 23.4schain gain +2.5 dBseam step 0.7 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(no disfluency, clear, didactic, formal) Tokens sollten ja schon am 17. bzw. 18. Dezember
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, formal; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 0.0/10; 4.4s, DE.
DE_-6F1Ss25DJQ_W000000 · in -19.4 dBFS · gain -0.7 dB · emolia-00037
(anger, disappointment, emotional numbness · some disfluency, average clarity, monologue, didactic) Ankommen sind sie aber nicht. Das Ganze hat sich einen Monat verzögert. Unbeam ist wie geplant Mitte Dezember tatsächlich live gegangen, aber hat dann erstmal einen Monat lang nur fast nur leere Blöcke produziert, um einfach zu testen, ob das Ganze funktioniert. Denn unsere
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as anger, disappointment, emotional numbness; style: monologue, didactic; good recording, no background noise; genuineness 2.7/6; vocal-burst blend 0.3/10; 19.2s, DE.
DE_-6F1Ss25DJQ_W000001 · in -19.2 dBFS · gain -0.8 dB · emolia-00037
Disappointment rising ↑identity +0.71 emotion 1 %   sad-Disappointment-S3-k2 · #4

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.00, 1.06. In the first clip the scorer found no Disappointment whatsoever (0.00); by the last it is at 1.06.

On the corpus-wide percentile scale those become 0.39, 0.95 — a total move of +0.56.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 28 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.020 before conversion and 0.727 after — it rose by 0.707. Neighbour-to-neighbour the worst pair went 0.020 → 0.727. (The earlier render, with segment 1 left raw, scores 0.642 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.557 in the original and +0.004 after conversion — 1 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.82 → 3.14 (+0.31) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.0615normalised 0.422 → 0.970identity cos to seg 1 0.020 → 0.727 +0.707identity cos neighbours 0.020 → 0.727d_b rescored +0.557 → +0.004d_a rescored +0.557 → +0.004d_a mined 1.062d_b mined 0.548min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-8Xv740j-8Qtotal 27.3schain gain +1.5 dBseam step 0.0 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a middle-aged feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(fear, anger, relief · clear, monologue, casual) Ich würde sagen, da ist auf jeden Fall Bedarf der Verbesserung, dass man sich nicht schämen muss, in diesem Beruf zu arbeiten. Es lohnt sich auf jeden Fall, weil der Schmutz stirbt nicht aus.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as fear, anger, relief; style: monologue, casual; good recording, quiet background; genuineness 2.0/6; vocal-burst blend 0.0/10; 10.1s, DE.
DE_-8Xv740j-8Q_W000000 · in -17.5 dBFS · gain -2.5 dB · emolia-00037
(anger, helplessness, shame · average clarity, monologue, casual) Was auch ja häufig ein Fehler ist, also ich kann das aus meiner eigenen Erfahrung mit Leistungen, mit Lieferanten, mit Bewerbern sagen, die die bei uns reinigen wollten, die Bürogebäude oder die Bürozimmer, dass die sich völlig falsch beworben haben. Entweder gab es kein Angebot oder man hat gleich gemerkt, dass die ja (low mumble) vielleicht nur schwarz arbeiten wollen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger, helplessness, shame; style: monologue, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 2.7/10; 17.5s, DE.
DE_-8Xv740j-8Q_W000017 · in -19.6 dBFS · gain -0.4 dB · emolia-00037
Disappointment rising ↑identity +0.16 emotion 76 %   sad-Disappointment-S3-k2 · #5

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.00, 1.28. In the first clip the scorer found no Disappointment whatsoever (0.00); by the last it is at 1.28.

On the corpus-wide percentile scale those become 0.39, 0.98 — a total move of +0.59.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 24 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.359 before conversion and 0.516 after — it rose by 0.157. Neighbour-to-neighbour the worst pair went 0.359 → 0.516. (The earlier render, with segment 1 left raw, scores 0.398 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.589 in the original and +0.447 after conversion — 76 % of the delta retained, which is most of it.

Quality. Mean predicted overall quality across the segments went 2.98 → 3.24 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.2832normalised 0.422 → 0.991identity cos to seg 1 0.359 → 0.516 +0.157identity cos neighbours 0.359 → 0.516d_b rescored +0.589 → +0.447d_a rescored +0.589 → +0.447d_a mined 1.283d_b mined 0.569min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-8c1BTgFP3ktotal 23.8schain gain +1.5 dBseam step 2.1 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(little disfluency, conversational, authoritative) Und jetzt der Wochendurchblick mit Florian Streibl. Liebe Zuschauerinnen und Zuschauer, willkommen zum Wochendurchblick.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational, authoritative; good recording, quiet background; genuineness 1.6/6; vocal-burst blend 0.3/10; 6.5s, DE.
DE_-8c1BTgFP3k_W000000 · in -18.5 dBFS · gain -1.5 dB · emolia-00247
(disappointment, impatience and irritability, anger · some disfluency, monologue, formal) Das waren Wahlleute, die ihre Hand weder für den Kandidaten von links noch von rechts heben wollten. Leute, die mit Amtshinthaber Frank-Walter Steinmeier durchaus zufrieden sind, die sich aber mehr Bürger mehr, mehr Ehrenamt und vor allem mehr Frauen in der Politik wünschen. Mit unserer Kandidatin.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, impatience and irritability, anger; style: monologue, formal; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.0/10; 17.5s, DE.
DE_-8c1BTgFP3k_W000005 · in -19.9 dBFS · gain -0.1 dB · emolia-00247
Disappointment rising ↑identity +0.00 emotion 27 %   sad-Disappointment-S3-k2 · #6

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.00, 1.29. In the first clip the scorer found no Disappointment whatsoever (0.00); by the last it is at 1.29.

On the corpus-wide percentile scale those become 0.39, 0.98 — a total move of +0.59.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 32 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.916 before conversion and 0.917 after — it rose by 0.000. Neighbour-to-neighbour the worst pair went 0.916 → 0.917. (The earlier render, with segment 1 left raw, scores 0.804 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.590 in the original and +0.162 after conversion — 27 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.93 → 3.10 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.2939normalised 0.422 → 0.991identity cos to seg 1 0.916 → 0.917 +0.000identity cos neighbours 0.916 → 0.917d_b rescored +0.590 → +0.162d_a rescored +0.590 → +0.162d_a mined 1.294d_b mined 0.569min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-HcL6XTJTVgtotal 31.3schain gain +2.2 dBseam step 0.7 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, brisk, energised, slightly relaxed
(elation, hope enthusiasm optimism, jealousy and envy · casual, dramatic) Hey Leute, ich bin's Conny und ich zeige euch, wie ihr in Roblox Studio euer eigenes Auto so erstellt, wie ihr es haben wollt. Und ich zeige euch auch, wie ihr einen Spawner für dieses Auto erstellt, damit ihr damit immer wieder neue Autos in die Map laden könnt. Aber zuallererst fangen wir an mit den Rädern. Dafür klicken wir hier oben auf Home, dann auf den File by Part,
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as elation, hope enthusiasm optimism, jealousy and envy; style: casual, dramatic; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 2.5/10; 18.3s, DE.
DE_-HcL6XTJTVg_W000000 · in -15.5 dBFS · gain -4.5 dB · emolia-00047
(disgust, sourness, disappointment · dramatic, casual) Und das 10 oder 100 ist egal. Ne große Zahl. Das ist dann die Stärke, wie sich das Lenkrad dreht. Und ihr seht, wenn wir das ausgewählt haben, das zeigt in diese Richtung. Das bedeutet, das ist falsch. Das ist falsch. Wir halten Alt gedrückt, nehmen dann dieses Attachment.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as disgust, sourness, disappointment; style: dramatic, casual; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 2.0/10; 13.3s, DE.
DE_-HcL6XTJTVg_W000012 · in -16.3 dBFS · gain -3.7 dB · emolia-00047
Disappointment rising ↑identity +0.34 emotion 93 %   sad-Disappointment-S3-k2 · #7

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.00, 1.11. In the first clip the scorer found no Disappointment whatsoever (0.00); by the last it is at 1.11.

On the corpus-wide percentile scale those become 0.39, 0.96 — a total move of +0.57.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 18 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.434 before conversion and 0.770 after — it rose by 0.336. Neighbour-to-neighbour the worst pair went 0.434 → 0.770. (The earlier render, with segment 1 left raw, scores 0.607 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.567 in the original and +0.528 after conversion — 93 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.85 → 3.00 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.1113normalised 0.422 → 0.977identity cos to seg 1 0.434 → 0.770 +0.336identity cos neighbours 0.434 → 0.770d_b rescored +0.567 → +0.528d_a rescored +0.567 → +0.528d_a mined 1.111d_b mined 0.555min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-Jk4ZBOnugMtotal 17.9schain gain -2.9 dBseam step 0.4 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady
(emotional numbness, helplessness, fear · average clarity, formal, storytelling) Egal was Sie machen, Ihre VR-Brille möchte einfach nicht über Ihre Seehilfe passen.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, helplessness, fear; style: formal, storytelling; good recording, quiet background; genuineness 2.2/6; vocal-burst blend 0.0/10; 4.7s, DE.
DE_-Jk4ZBOnugM_W000000 · in -16.3 dBFS · gain -3.7 dB · emolia-00147
(helplessness, disappointment, fatigue exhaustion · somewhat unclear, monologue, casual) Aber trotzdem so stabil, dass sie nicht kaputt gehen. Also, ich glaub nicht, da muss man schon mit Gewalt dran (low mumble) äh reißen, dass da irgendwas abreißt. Und das sieht schon echt stabil aus. Und, (low mumble) äh, im gepflegten Umgang damit, äh, sollte da echt nichts passieren.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness, disappointment, fatigue exhaustion; style: monologue, casual; average recording, no background noise; genuineness 4.6/6; vocal-burst blend 3.5/10; 13.5s, DE.
DE_-Jk4ZBOnugM_W000018 · in -18.5 dBFS · gain -1.5 dB · emolia-00147
Disappointment rising ↑identity −0.00 emotion 96 %   sad-Disappointment-S3-k2 · #8

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.00, 1.17. In the first clip the scorer found no Disappointment whatsoever (0.00); by the last it is at 1.17.

On the corpus-wide percentile scale those become 0.39, 0.97 — a total move of +0.58.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 20 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.907 before conversion and 0.906 after — it fell by 0.001. Neighbour-to-neighbour the worst pair went 0.907 → 0.906. (The earlier render, with segment 1 left raw, scores 0.814 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.577 in the original and +0.553 after conversion — 96 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.84 → 3.21 (+0.37) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.1699normalised 0.422 → 0.983identity cos to seg 1 0.907 → 0.906 -0.001identity cos neighbours 0.907 → 0.906d_b rescored +0.577 → +0.553d_a rescored +0.577 → +0.553d_a mined 1.170d_b mined 0.561min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-NcjdeCDcbYtotal 19.3schain gain +3.7 dBseam step 0.2 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normally alert, slightly relaxed, fairly steady
(disgust · normal-paced, little disfluency, clear, monologue) Wir sind urbex.le direkt aus Leipzig und in dieser Folge zeigen wir euch eine ehemalige Gaststätte und Wohnhaus. Kleiner Hinweis, zum Schutz des Gebäudes werde ich den Gaststättennamen nicht nennen.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust; style: monologue, didactic; good recording, quiet background; genuineness 0.4/6; vocal-burst blend 0.5/10; 11.2s, DE.
DE_-NcjdeCDcbY_W000001 · in -21.0 dBFS · gain +1.0 dB · emolia-00182
(disappointment, confusion · brisk, some disfluency, average clarity, casual) Das es sich hierbei um eine Gaststätte handeln sollte, konnten wir vor Ort nicht bejahen, denn davon war rein gar nichts mehr zu erkennen, weder von außen noch von innen.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as disappointment, confusion; style: casual, dramatic; good recording, no background noise; genuineness 3.1/6; vocal-burst blend 0.5/10; 8.3s, DE.
DE_-NcjdeCDcbY_W000002 · in -22.0 dBFS · gain +2.0 dB · emolia-00182
Disappointment rising ↑identity +0.01 emotion 85 %   sad-Disappointment-S3-k2 · #9

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.00, 1.20. In the first clip the scorer found no Disappointment whatsoever (0.00); by the last it is at 1.20.

On the corpus-wide percentile scale those become 0.39, 0.98 — a total move of +0.58.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 16 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.710 before conversion and 0.723 after — it rose by 0.012. Neighbour-to-neighbour the worst pair went 0.710 → 0.723. (The earlier render, with segment 1 left raw, scores 0.606 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.581 in the original and +0.496 after conversion — 85 % of the delta retained, which is most of it.

Quality. Mean predicted overall quality across the segments went 2.61 → 2.94 (+0.33) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.2031normalised 0.422 → 0.986identity cos to seg 1 0.710 → 0.723 +0.012identity cos neighbours 0.710 → 0.723d_b rescored +0.581 → +0.496d_a rescored +0.581 → +0.496d_a mined 1.203d_b mined 0.564min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-OLjkEFs1wktotal 15.7schain gain +1.4 dBseam step 0.5 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, balanced body, average recording, normally alert, slightly relaxed, wide pitch range
(thankfulness gratitude, affection · measured, fairly steady, frequent disfluency, authoritative) Frau Präsidentin, meine sehr geehrten Damen und Herren! Ich gebe den Bericht.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, affection; style: authoritative; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 0.0/10; 5.0s, DE.
DE_-OLjkEFs1wk_W000000 · in -19.3 dBFS · gain -0.7 dB · emolia-00247
(disappointment, bitterness, anger · normal-paced, moderately variable, some disfluency, authoritative) Der Älteste hat den Dringlichen Gesetzentwurf in seiner 27. Sitzung am 14. Juni 2016 verraten und die unter A wiedergegebene Beschlussempfehlung an das Plenum ausgesprochen, die wie vorher gelautet.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as disappointment, bitterness, anger; style: authoritative, monologue; average recording, some background noise; genuineness 2.2/6; vocal-burst blend 1.5/10; 11.0s, DE.
DE_-OLjkEFs1wk_W000002 · in -18.7 dBFS · gain -1.3 dB · emolia-00247
Disappointment rising ↑identity −0.00 emotion 100 %   sad-Disappointment-S3-k2 · #10

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.00, 1.35. In the first clip the scorer found no Disappointment whatsoever (0.00); by the last it is at 1.35.

On the corpus-wide percentile scale those become 0.39, 0.99 — a total move of +0.59.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 36 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.902 before conversion and 0.901 after — it fell by 0.001. Neighbour-to-neighbour the worst pair went 0.902 → 0.901. (The earlier render, with segment 1 left raw, scores 0.840 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.594 in the original and +0.596 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.94 → 3.25 (+0.31) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.3506normalised 0.422 → 0.993identity cos to seg 1 0.902 → 0.901 -0.001identity cos neighbours 0.902 → 0.901d_b rescored +0.594 → +0.596d_a rescored +0.594 → +0.596d_a mined 1.351d_b mined 0.572min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-U2hZsyyyRutotal 35.2schain gain +5.4 dBseam step 0.6 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult somewhat feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, subdued, slightly relaxed
(malevolence malice · measured, frequent disfluency, clear, didactic) Heute möchte ich ein Seifenblasenpapier machen. Das heißt, ich puste Seifenblasen, bunte Seifenblasen auf ein Papier, die darauf platzen und dort ihre farbige Spur hinterlassen.
full caption & clip details
A young adult somewhat feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, slightly guarded; reads as malevolence malice; style: didactic, whispered; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.5/10; 15.6s, DE.
DE_-U2hZsyyyRu_W000000 · in -22.5 dBFS · gain +2.5 dB · emolia-00147
(disappointment, shame, fatigue exhaustion · normal-paced, some disfluency, average clarity, casual) Hier habe ich, (low mumble) äh, ein Lesezeichen gemacht und ich verwende dann gerne einen Eckenabrunder. Die gibt's auch zu kaufen. Das sind so kleine Geräte. Also, ich hab hier so eins. Das hat drei verschiedene Einstellungen und ist dann, also das veredelt dann diese, diese Arbeiten einfach nochmal.
full caption & clip details
A young adult somewhat feminine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as disappointment, shame, fatigue exhaustion; style: casual, monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 0.0/10; 19.9s, DE.
DE_-U2hZsyyyRu_W000016 · in -21.0 dBFS · gain +1.0 dB · emolia-00147
Disappointment rising ↑identity +0.10 emotion REVERSED   sad-Disappointment-S3-k2 · #11

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.00, 1.22. In the first clip the scorer found no Disappointment whatsoever (0.00); by the last it is at 1.22.

On the corpus-wide percentile scale those become 0.39, 0.98 — a total move of +0.58.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 19 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.802 before conversion and 0.906 after — it rose by 0.104. Neighbour-to-neighbour the worst pair went 0.802 → 0.906. (The earlier render, with segment 1 left raw, scores 0.776 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.583 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.99 → 3.24 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.2207normalised 0.422 → 0.987identity cos to seg 1 0.802 → 0.906 +0.104identity cos neighbours 0.802 → 0.906d_b rescored +0.583 → +0.000d_a rescored +0.583 → +0.000d_a mined 1.221d_b mined 0.565min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-WhG-j-UudUtotal 19.0schain gain +2.9 dBseam step 0.8 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normally alert, slightly relaxed, fairly steady
(measured, little disfluency, authoritative, didactic) Hello. Herzlich willkommen aus der Quantum Storm Star Wars Collection. Mein Name ist Deniz und ich möchte euch heute den
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, didactic; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 0.1/10; 6.7s, DE.
DE_-WhG-j-UudU_W000000 · in -19.3 dBFS · gain -0.7 dB · emolia-00238
(bitterness, emotional numbness, disappointment · normal-paced, some disfluency, monologue, casual) Es wurde von der Firma Kenner schon 1979 einmal herausgebracht, aber damals wurde es noch in das sogenannte Expanded Universe verortet, was ja jetzt nicht mehr ganz richtig ist, da die Serie ja zum Kanon gehört.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as bitterness, emotional numbness, disappointment; style: monologue, casual; good recording, quiet background; genuineness 3.0/6; vocal-burst blend 1.1/10; 12.6s, DE.
DE_-WhG-j-UudU_W000008 · in -19.6 dBFS · gain -0.4 dB · emolia-00238
Disappointment rising ↑identity +0.03 emotion 93 %   sad-Disappointment-S3-k2 · #12

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.00, 1.18. In the first clip the scorer found no Disappointment whatsoever (0.00); by the last it is at 1.18.

On the corpus-wide percentile scale those become 0.39, 0.97 — a total move of +0.58.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 32 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.872 before conversion and 0.900 after — it rose by 0.027. Neighbour-to-neighbour the worst pair went 0.872 → 0.900. (The earlier render, with segment 1 left raw, scores 0.879 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.578 in the original and +0.537 after conversion — 93 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.83 → 3.22 (+0.39) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.1797normalised 0.422 → 0.984identity cos to seg 1 0.872 → 0.900 +0.027identity cos neighbours 0.872 → 0.900d_b rescored +0.578 → +0.537d_a rescored +0.578 → +0.537d_a mined 1.180d_b mined 0.562min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-YzjWuJuZL8total 31.7schain gain +2.2 dBseam step 0.6 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(hope enthusiasm optimism, elation, interest · casual, conversational) Hallo in die Runde. Ich freue mich sehr, dass ihr dabei seid, mit mir heute wieder über Skat sprechen wollt. Da könnt ihr es auch schon sehen, mir ist demnetzt tatsächlich mal wieder eine ganz interessante Partie über den Weg gelaufen, online, wie ihr seht.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, elation, interest; style: casual, conversational; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 3.8/10; 13.1s, DE.
DE_-YzjWuJuZL8_W000000 · in -17.1 dBFS · gain -2.9 dB · emolia-00238
(pain, doubt, bitterness · monologue, casual) Von mir die Pik Dame, Alleinspieler Sticht. Der Stich sieht für den Alleinspieler sicher maximal harmlos aus. Wenn da so großzügig Asse angeboten werden, sitzt normalerweise die Karte relativ glatt. Entsprechend ist er jetzt sicher ein wenig überrascht, dass hier die Trümpfe zu viert dagegen stehen.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain, doubt, bitterness; style: monologue, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 0.3/10; 18.9s, DE.
DE_-YzjWuJuZL8_W000008 · in -18.7 dBFS · gain -1.3 dB · emolia-00238
Disappointment rising ↑identity +0.04 emotion 83 %   sad-Disappointment-S3-k2 · #13

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.00, 1.24. In the first clip the scorer found no Disappointment whatsoever (0.00); by the last it is at 1.24.

On the corpus-wide percentile scale those become 0.39, 0.98 — a total move of +0.59.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 26 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.851 before conversion and 0.890 after — it rose by 0.039. Neighbour-to-neighbour the worst pair went 0.851 → 0.890. (The earlier render, with segment 1 left raw, scores 0.776 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.586 in the original and +0.485 after conversion — 83 % of the delta retained, which is most of it.

Quality. Mean predicted overall quality across the segments went 3.03 → 3.25 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.2451normalised 0.422 → 0.988identity cos to seg 1 0.851 → 0.890 +0.039identity cos neighbours 0.851 → 0.890d_b rescored +0.586 → +0.485d_a rescored +0.586 → +0.485d_a mined 1.245d_b mined 0.567min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-bnJC1eLO5Itotal 25.6schain gain +2.8 dBseam step 1.4 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, some disfluency
(slightly relaxed, fairly steady, moderate pitch range, storytelling) Wir machen eine Geisbergrunde heute miteinander. Man kann um den ganzen Geisberg rum spazieren. Und das Video, das werden wir heute rund um den Geisberg machen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: storytelling, monologue; good recording, quiet background; genuineness 1.9/6; vocal-burst blend 1.0/10; 9.6s, DE.
DE_-bnJC1eLO5I_W000000 · in -18.2 dBFS · gain -1.8 dB · emolia-00113
(impatience and irritability, disappointment, anger · neutral tension, moderately variable, wide pitch range, didactic) Frage, nutzt es jetzt etwas, wenn wir sagen, nein, es darf jetzt nicht sein, da darf jetzt nicht umfallen, das können wir jetzt nicht brauchen und da damit haben wir jetzt nicht gerechnet, nein, es geht jetzt nicht. Oder die Türe ist jetzt zugongen, nein, es geht jetzt nicht, das können wir jetzt nicht haben. Der Baum liegt, wenn wir sehen, gäh.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, disappointment, anger; style: didactic, monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 3.7/10; 16.3s, DE.
DE_-bnJC1eLO5I_W000005 · in -17.6 dBFS · gain -2.4 dB · emolia-00113
Disappointment rising ↑identity +0.00 emotion 96 %   sad-Disappointment-S3-k2 · #14

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.00, 1.33. In the first clip the scorer found no Disappointment whatsoever (0.00); by the last it is at 1.33.

On the corpus-wide percentile scale those become 0.39, 0.99 — a total move of +0.59.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 46 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.897 before conversion and 0.899 after — it rose by 0.002. Neighbour-to-neighbour the worst pair went 0.897 → 0.899. (The earlier render, with segment 1 left raw, scores 0.787 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.593 in the original and +0.570 after conversion — 96 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.82 → 3.10 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.332normalised 0.422 → 0.993identity cos to seg 1 0.897 → 0.899 +0.002identity cos neighbours 0.897 → 0.899d_b rescored +0.593 → +0.570d_a rescored +0.593 → +0.570d_a mined 1.332d_b mined 0.571min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-e7x_OIHd0gtotal 45.5schain gain +1.8 dBseam step 0.4 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normally alert, slightly relaxed, fairly steady
(disgust, concentration, infatuation · brisk, clear, newsreading, formal) Am Freitag startet der CDU-Bundesparteitag in Leipzig. Der CDU-Bundesparteitag ist das höchste beschlussfassende Gremium der CDU Deutschlands. Was sich kompliziert anhört, ist eigentlich ganz einfach. Dort werden die wichtigsten inhaltlichen und personellen Entscheidungen gekriegt.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as disgust, concentration, infatuation; style: newsreading, formal; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.1/10; 15.6s, DE.
DE_-e7x_OIHd0g_W000000 · in -19.9 dBFS · gain -0.1 dB · emolia-00247
(bitterness, disappointment, triumph · normal-paced, average clarity, monologue, authoritative) Aus ganz Deutschland kommen rund 1000 Delegierte nach Leipzig.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as bitterness, disappointment, triumph; style: monologue, authoritative; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 1.9/10; 30.0s, DE.
DE_-e7x_OIHd0g_W000001 · in -18.8 dBFS · gain -1.2 dB · emolia-00247
Disappointment rising ↑identity +0.08 emotion 339 %   sad-Disappointment-S3-k2 · #15

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.01, 1.05. In the first clip the scorer found no Disappointment whatsoever (0.01); by the last it is at 1.05.

On the corpus-wide percentile scale those become 0.79, 0.95 — a total move of +0.16.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 14 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.761 before conversion and 0.840 after — it rose by 0.078. Neighbour-to-neighbour the worst pair went 0.761 → 0.840. (The earlier render, with segment 1 left raw, scores 0.803 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.158 in the original and +0.535 after conversion — 339 % of the delta retained, i.e. the move came out slightly larger after conversion than before.

Quality. Mean predicted overall quality across the segments went 2.56 → 2.94 (+0.38) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0092 → 1.0537normalised 0.845 → 0.968identity cos to seg 1 0.761 → 0.840 +0.078identity cos neighbours 0.761 → 0.840d_b rescored +0.158 → +0.535d_a rescored +0.158 → +0.535d_a mined 1.044d_b mined 0.123min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-f6GmCGhSBototal 13.2schain gain +0.8 dBseam step 4.4 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, quiet background, normal-paced, normally alert
(sourness, shame, thankfulness gratitude · conversational, casual) Ich will euch heute zeigen, was ich wirklich innerhalb, wenn man Mutter ist, dann hat man wenig Zeit und was ich wirklich so
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, neutral stance, neutral openness; reads as sourness, shame, thankfulness gratitude; style: conversational, casual; good recording, quiet background; genuineness 4.7/6; vocal-burst blend 2.4/10; 5.9s, DE.
DE_-f6GmCGhSBo_W000000 · in -19.7 dBFS · gain -0.3 dB · emolia-00224
(disgust, sourness, impatience and irritability · dramatic, didactic) macht das meistens, guckt noch nicht mehr in den Spiegel, also das ist alles schon so routiniert, dass ich das einfach so schnell auftrage. Na, sehe ich schon fresher aus?
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as disgust, sourness, impatience and irritability; style: dramatic, didactic; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.0/10; 7.6s, DE.
DE_-f6GmCGhSBo_W000003 · in -23.3 dBFS · gain +3.3 dB · emolia-00224
Disappointment rising ↑identity +0.97 emotion 99 %   sad-Disappointment-S3-k2 · #16

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.00, 1.38. In the first clip the scorer found no Disappointment whatsoever (0.00); by the last it is at 1.38.

On the corpus-wide percentile scale those become 0.39, 0.99 — a total move of +0.60.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 34 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.024 before conversion and 0.942 after — it rose by 0.966. Neighbour-to-neighbour the worst pair went -0.024 → 0.942. (The earlier render, with segment 1 left raw, scores 0.792 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.595 in the original and +0.587 after conversion — 99 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 3.14 → 3.33 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.3799normalised 0.422 → 0.994identity cos to seg 1 -0.024 → 0.942 +0.966identity cos neighbours -0.024 → 0.942d_b rescored +0.595 → +0.587d_a rescored +0.595 → +0.587d_a mined 1.380d_b mined 0.573min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-ffmw_U0FWYtotal 33.9schain gain +2.2 dBseam step 0.1 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normal-paced, normally alert, slightly relaxed
(concentration, pride · formal, newsreading) Mehr Tempo auf dem Weg zur Klimaneutralität, das fordert die Organisation für wirtschaftliche Entwicklung und Zusammenarbeit in ihrem neuesten Umweltprüfbericht von der Bundesregierung. Ein Appell der OECD lautet deshalb mehr E-Mobilität und mehr Verkehr auf die Schiene.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, pride; style: formal, newsreading; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.0/10; 16.4s, DE.
DE_-ffmw_U0FWY_W000000 · in -16.3 dBFS · gain -3.7 dB · emolia-00147
(disappointment, concentration, anger · formal, monologue) Die Wirtschaft wachse, obwohl die Emissionen sinken, hebt die OECD hervor. Aber noch immer kommen rund drei Viertel des Energieaufkommens aus frustilen Quellen. Nötig sei der Ausbau von E-Mobilität und Schiene, Aufforstung von Wäldern, Renatuierung der Moore und der Schutz vor Auswirkungen des Klimawandels.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, concentration, anger; style: formal, monologue; average recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 17.7s, DE.
DE_-ffmw_U0FWY_W000006 · in -14.4 dBFS · gain -5.6 dB · emolia-00147
Disappointment rising ↑identity +0.04 emotion 90 %   sad-Disappointment-S3-k2 · #17

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.00, 1.18. In the first clip the scorer found no Disappointment whatsoever (0.00); by the last it is at 1.18.

On the corpus-wide percentile scale those become 0.39, 0.97 — a total move of +0.58.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 18 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.735 before conversion and 0.771 after — it rose by 0.036. Neighbour-to-neighbour the worst pair went 0.735 → 0.771. (The earlier render, with segment 1 left raw, scores 0.750 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.579 in the original and +0.521 after conversion — 90 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.76 → 2.99 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.1816normalised 0.422 → 0.984identity cos to seg 1 0.735 → 0.771 +0.036identity cos neighbours 0.735 → 0.771d_b rescored +0.579 → +0.521d_a rescored +0.579 → +0.521d_a mined 1.182d_b mined 0.562min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-iFXnOKFbJktotal 18.0schain gain +1.4 dBseam step 0.6 dBcrossfades 150 ms
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(pride, disgust · average clarity, conversational, casual) Dann ist der Fahrradweg hier nicht mehr (ahem) zu befahren. Also der steht meistens so bis knietief unter Wasser.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, disgust; style: conversational, casual; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 0.0/10; 5.3s, DE.
DE_-iFXnOKFbJk_W000001 · in -24.4 dBFS · gain +4.4 dB · emolia-00095
(disappointment, distress, bitterness · somewhat unclear, casual, conversational) Und, (low mumble) ähm, das haben wir dann einfach gemacht, wie wenn wir so einen Hochwassersensor angebracht haben. Ich spring nochmal zurück, weil das das falsche Bild war. Man sieht da oben an der Spitze oder an der, in diesem dunklen Bereich sieht man den Hochwassersensor. Das ist so eine kleine,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, distress, bitterness; style: casual, conversational; average recording, quiet background; genuineness 5.6/6; vocal-burst blend 1.5/10; 12.9s, DE.
DE_-iFXnOKFbJk_W000004 · in -23.4 dBFS · gain +3.4 dB · emolia-00095
Disappointment rising ↑identity +0.36 emotion 98 %   sad-Disappointment-S3-k2 · #18

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is -0.00, 1.10. In the first clip the scorer found no Disappointment whatsoever (-0.00); by the last it is at 1.10.

On the corpus-wide percentile scale those become 0.39, 0.96 — a total move of +0.57.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 22 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.281 before conversion and 0.642 after — it rose by 0.362. Neighbour-to-neighbour the worst pair went 0.281 → 0.642. (The earlier render, with segment 1 left raw, scores 0.542 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.960 in the original and +0.940 after conversion — 98 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.74 → 3.13 (+0.38) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.1025normalised 0.422 → 0.976identity cos to seg 1 0.281 → 0.642 +0.362identity cos neighbours 0.281 → 0.642d_b rescored +0.960 → +0.940d_a rescored +0.960 → +0.940d_a mined 1.103d_b mined 0.554min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-xg8lm69_K8total 22.0schain gain +3.2 dBseam step 3.5 dBcrossfades 100 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-bright, fairly smooth, balanced body, average recording, normal-paced, energised, wide pitch range
(slightly relaxed, fairly steady, frequent disfluency, authoritative) Und ich rufe auf den Tagesordnungspunkt 11.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very clear, frequent disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, formal; average recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.3/10; 3.2s, DE.
DE_-xg8lm69_K8_W000000 · in -18.2 dBFS · gain -1.8 dB · emolia-00224
(anger, sourness, jealousy and envy · neutral tension, moderately variable, some disfluency, ranting) (ahem) gibt es nur noch eine einzige staatliche Hochschule, die mehr als 20 % Wahlbeteiligung bei Studierenden hat. Und die Tendenz ist ansonsten überall fallend. Meine Damen und Herren, nicht Wahl oder nicht zur Wahl zu gehen, ist natürlich das Recht von jedem. Aber ich finde, das darf uns nicht gleichgültig lassen.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as anger, sourness, jealousy and envy; style: ranting, cartoonish; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 2.1/10; 18.9s, DE.
DE_-xg8lm69_K8_W000004 · in -24.1 dBFS · gain +4.1 dB · emolia-00224
Disappointment rising ↑identity −0.04 emotion REVERSED   sad-Disappointment-S3-k2 · #19

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is 0.00, 1.08. In the first clip the scorer found no Disappointment whatsoever (0.00); by the last it is at 1.08.

On the corpus-wide percentile scale those become 0.39, 0.96 — a total move of +0.56.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 25 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.746 before conversion and 0.710 after — it fell by 0.037. Neighbour-to-neighbour the worst pair went 0.746 → 0.710. (The earlier render, with segment 1 left raw, scores 0.731 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.561 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.60 → 2.97 (+0.38) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.0801normalised 0.422 → 0.973identity cos to seg 1 0.746 → 0.710 -0.037identity cos neighbours 0.746 → 0.710d_b rescored +0.561 → +0.000d_a rescored +0.561 → +0.000d_a mined 1.080d_b mined 0.551min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-xpBDPzvHIItotal 24.7schain gain +2.9 dBseam step 1.9 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, quiet background, normally alert, slightly relaxed
(doubt · measured, fairly steady, clear, formal) Du kennst bestimmt so die Situation. Du hast schon lange irgendwas nicht mehr gegessen.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as doubt; style: formal, didactic; good recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.0/10; 5.3s, DE.
DE_-xpBDPzvHII_W000000 · in -18.9 dBFS · gain -1.1 dB · emolia-00182
(disappointment, awe, jealousy and envy · normal-paced, moderately variable, average clarity, casual) Können wir mit allem machen? Wir können auch Joghurt nehmen, ja? Du gehst, du stehst vorm Regal im Supermarkt, vorm Joghurtregal und guckst, welche Verpackungen. Ist es Glas? Ist es das Glas, was einfach nur rund ist? Oder ist es das Glas, was ausschaut wie so ein kleines Fässchen? Oder ist es das Glas?
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as disappointment, awe, jealousy and envy; style: casual, conversational; good recording, quiet background; genuineness 3.0/6; vocal-burst blend 1.1/10; 19.6s, DE.
DE_-xpBDPzvHII_W000005 · in -17.9 dBFS · gain -2.1 dB · emolia-00182
Disappointment rising ↑identity +0.09 emotion 95 %   sad-Disappointment-S3-k2 · #20

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Disappointment.

The raw scorer output across the chain is -0.00, 1.31. In the first clip the scorer found no Disappointment whatsoever (-0.00); by the last it is at 1.31.

On the corpus-wide percentile scale those become 0.39, 0.99 — a total move of +0.59.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Disappointment, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Disappointment sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 21 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.775 before conversion and 0.862 after — it rose by 0.087. Neighbour-to-neighbour the worst pair went 0.775 → 0.862. (The earlier render, with segment 1 left raw, scores 0.776 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.986 in the original and +0.940 after conversion — 95 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 3.00 → 3.20 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.3096normalised 0.422 → 0.992identity cos to seg 1 0.775 → 0.862 +0.087identity cos neighbours 0.775 → 0.862d_b rescored +0.986 → +0.940d_a rescored +0.986 → +0.940d_a mined 1.310d_b mined 0.570min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0068rhxjjg8total 21.1schain gain +2.1 dBseam step 0.2 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, slightly dark, balanced body, average recording, quiet background, slightly relaxed, steady, somewhat unclear
(measured, normally alert, some disfluency, didactic) Die zehnte Etappe der Tour de France 2022 startete in Morsin, in Haute-Savoyen.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 0.1/10; 6.9s, DE.
DE_0068rhxjjg8_W000000 · in -20.8 dBFS · gain +0.8 dB · emolia-00215
(disappointment, bitterness, anger · slow, very low-energy, frequent disfluency, monologue) Die Werbeparade fuhr natürlich auch dieses Mal wieder durch, wie wir sie schon im letzten Video gezeigt haben. Keine Sorge, wir zeigen sie nicht noch einmal, nur die Ergebnisse, die sie für uns gebracht hat.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, bitterness, anger; style: monologue, whispered; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.7/10; 14.4s, DE.
DE_0068rhxjjg8_W000007 · in -22.4 dBFS · gain +2.4 dB · emolia-00215