rescue rule sad-Distress-S3-k2 — voice-corrected

Distress under rescue rule S3, k=2. generalisation check, gap 0.453. 90.5 % of clips score at or below zero on this emotion and the largest gap on its normalised axis is 0.453 (WIDER than the 0.25 step cap). This rule found 7,875 chains over 40,000 tracks; the strict rule found 0 at k=3.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_sad-Distress-S3-k2.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
These are not strict-rule trajectories. They come from a deliberately looser rule, built to recover examples on an emotion the strict rule cannot reach. What rule S3 changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'. What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained. Full explanation →
20chains converted
20segments re-voiced
0.718 → 0.795median worst-to-anchor identity cosine
95 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Distress rising ↑identity −0.03 emotion 98 %   sad-Distress-S3-k2 · #1

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.19. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.19.

On the corpus-wide percentile scale those become 0.44, 0.99 — a total move of +0.54.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 25 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.911 before conversion and 0.880 after — it fell by 0.031. Neighbour-to-neighbour the worst pair went 0.911 → 0.880. (The earlier render, with segment 1 left raw, scores 0.865 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.545 in the original and +0.536 after conversion — 98 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.87 → 3.23 (+0.36) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.1865normalised 0.453 → 0.988identity cos to seg 1 0.911 → 0.880 -0.031identity cos neighbours 0.911 → 0.880d_b rescored +0.545 → +0.536d_a rescored +0.545 → +0.536d_a mined 1.187d_b mined 0.535min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-YzjWuJuZL8total 24.9schain gain +2.1 dBseam step 0.1 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(hope enthusiasm optimism, elation, interest · casual, conversational) Hallo in die Runde. Ich freue mich sehr, dass ihr dabei seid, mit mir heute wieder über Skat sprechen wollt. Da könnt ihr es auch schon sehen, mir ist demnetzt tatsächlich mal wieder eine ganz interessante Partie über den Weg gelaufen, online, wie ihr seht.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, elation, interest; style: casual, conversational; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 3.8/10; 13.1s, DE.
DE_-YzjWuJuZL8_W000000 · in -17.1 dBFS · gain -2.9 dB · emolia-00238
(doubt, distress, emotional numbness · monologue, casual) sowieso nicht mehr verteidigen könnt gegen keine Verteilung. Diese Karte ist de facto schon tot. Das haben wir eben im Spielverlauf auch gesehen. Kurz vor Schluss musste die Kreuz 10 sowieso angeboten werden.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, distress, emotional numbness; style: monologue, casual; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 0.0/10; 12.1s, DE.
DE_-YzjWuJuZL8_W000019 · in -17.8 dBFS · gain -2.2 dB · emolia-00238
Distress rising ↑identity +0.08 emotion 612 %   sad-Distress-S3-k2 · #2

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.02, 1.02. In the first clip the scorer found no Distress whatsoever (0.02); by the last it is at 1.02.

On the corpus-wide percentile scale those become 0.89, 0.98 — a total move of +0.09.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 18 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.695 before conversion and 0.780 after — it rose by 0.084. Neighbour-to-neighbour the worst pair went 0.695 → 0.780. (The earlier render, with segment 1 left raw, scores 0.689 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.086 in the original and +0.525 after conversion — 612 % of the delta retained, i.e. the move came out slightly larger after conversion than before.

Quality. Mean predicted overall quality across the segments went 2.75 → 3.07 (+0.33) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0164 → 1.0244normalised 0.913 → 0.982identity cos to seg 1 0.695 → 0.780 +0.084identity cos neighbours 0.695 → 0.780d_b rescored +0.086 → +0.525d_a rescored +0.086 → +0.525d_a mined 1.008d_b mined 0.069min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-f6GmCGhSBototal 18.1schain gain +3.8 dBseam step 0.2 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-bright, fairly smooth, normal-paced, normally alert, neutral tension, moderately variable, light breath
(sourness, shame, thankfulness gratitude · some disfluency, average clarity, wide pitch range, conversational) Ich will euch heute zeigen, was ich wirklich innerhalb, wenn man Mutter ist, dann hat man wenig Zeit und was ich wirklich so
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, neutral stance, neutral openness; reads as sourness, shame, thankfulness gratitude; style: conversational, casual; good recording, quiet background; genuineness 4.7/6; vocal-burst blend 2.4/10; 5.9s, DE.
DE_-f6GmCGhSBo_W000000 · in -19.7 dBFS · gain -0.3 dB · emolia-00224
(jealousy and envy, confusion, relief · frequent disfluency, somewhat unclear, fairly narrow pitch, casual) Sorry, love. Also, ihr habt bestimmt auch echt geile Produkte. Jetzt fragt ihr euch wahrscheinlich, wo hast du die denn her? Gibt's die jetzt nur in der Türkei? Nee, die hat mir eine liebe Kohle hingestellt. Irgendwie verklebt, das sieht alles so klumpig aus. Also,
full caption & clip details
A young adult somewhat masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as jealousy and envy, confusion, relief; style: casual, conversational; below-average recording, some background noise; genuineness 4.4/6; vocal-burst blend 7.6/10; 12.4s, DE.
DE_-f6GmCGhSBo_W000010 · in -22.3 dBFS · gain +2.3 dB · emolia-00224
Distress rising ↑identity +0.18 emotion 97 %   sad-Distress-S3-k2 · #3

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.47. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.47.

On the corpus-wide percentile scale those become 0.44, 0.99 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 33 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.304 before conversion and 0.488 after — it rose by 0.185. Neighbour-to-neighbour the worst pair went 0.304 → 0.488. (The earlier render, with segment 1 left raw, scores 0.392 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.552 in the original and +0.533 after conversion — 97 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.68 → 3.06 (+0.37) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.4727normalised 0.453 → 0.995identity cos to seg 1 0.304 → 0.488 +0.185identity cos neighbours 0.304 → 0.488d_b rescored +0.552 → +0.533d_a rescored +0.552 → +0.533d_a mined 1.473d_b mined 0.542min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-xg8lm69_K8total 33.1schain gain +1.7 dBseam step 1.5 dBcrossfades 100 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-bright, fairly smooth, balanced body, average recording, energised, light breath
(normal-paced, slightly relaxed, fairly steady, authoritative) Und ich rufe auf den Tagesordnungspunkt 11.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very clear, frequent disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, formal; average recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.3/10; 3.2s, DE.
DE_-xg8lm69_K8_W000000 · in -18.2 dBFS · gain -1.8 dB · emolia-00224
(disappointment, distress, impatience and irritability · brisk, neutral tension, moderately variable, monologue) Und da würde ich sagen, liegt es nicht nur daran, dass das Wahlprozedere jetzt irgendwie kompliziert wäre. Ich meine, die Studierenden bekommen die Unterlagen geschickt. Es gibt tagelang Zeit, seine Stimme abzugeben. Ich denke, wir müssen auch überlegen, ob es was damit zu tun hat, (ahem) dass die Entscheidungskompetenzen in der Verfasst Studierendenschaft nicht gerade sehr ausgeprägt sind, um es vorsichtig zu sagen. Also, den ganzen Autonomieprozess, den es gab, sind ja Kompetenzen vom Ministerium an die Hochschulen (ahem) (ahem) verlagert worden, aber eben dort vor allem an die Präsidien und an die Hochschulräte. Und ich finde, man muss vielleicht
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as disappointment, distress, impatience and irritability; style: monologue, dramatic; average recording, some background noise; genuineness 2.1/6; vocal-burst blend 2.7/10; 30.0s, DE.
DE_-xg8lm69_K8_W000052 · in -24.0 dBFS · gain +4.0 dB · emolia-00224
Distress rising ↑identity +0.08 emotion 101 %   sad-Distress-S3-k2 · #4

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.02. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.02.

On the corpus-wide percentile scale those become 0.44, 0.98 — a total move of +0.54.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 17 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.760 before conversion and 0.841 after — it rose by 0.081. Neighbour-to-neighbour the worst pair went 0.760 → 0.841. (The earlier render, with segment 1 left raw, scores 0.841 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.537 in the original and +0.540 after conversion — 101 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.75 → 3.21 (+0.45) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.0205normalised 0.453 → 0.982identity cos to seg 1 0.760 → 0.841 +0.081identity cos neighbours 0.760 → 0.841d_b rescored +0.537 → +0.540d_a rescored +0.537 → +0.540d_a mined 1.020d_b mined 0.529min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-xpBDPzvHIItotal 17.1schain gain +2.3 dBseam step 1.1 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, measured, normally alert, slightly relaxed
(doubt · fairly steady, formal, didactic) Du kennst bestimmt so die Situation. Du hast schon lange irgendwas nicht mehr gegessen.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as doubt; style: formal, didactic; good recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.0/10; 5.3s, DE.
DE_-xpBDPzvHII_W000000 · in -18.9 dBFS · gain -1.1 dB · emolia-00182
(jealousy and envy, distress, impatience and irritability · moderately variable, didactic, monologue) Und da kannst du dich doch nur entscheiden von deinem Gefühl her, wie im Supermarkt, wenn du dir ein Müsli kaufst, welche Packen gefällt mir am besten. Und das holst du dir dann und alles andere.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as jealousy and envy, distress, impatience and irritability; style: didactic, monologue; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.0/10; 12.0s, DE.
DE_-xpBDPzvHII_W000019 · in -17.9 dBFS · gain -2.1 dB · emolia-00182
Distress rising ↑identity −0.13 emotion REVERSED   sad-Distress-S3-k2 · #5

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.24. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.24.

On the corpus-wide percentile scale those become 0.44, 0.99 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 14 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.741 before conversion and 0.616 after — it fell by 0.125. Neighbour-to-neighbour the worst pair went 0.741 → 0.616. (The earlier render, with segment 1 left raw, scores 0.369 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Distress moved +0.546 in the original and -0.043 after conversion — it changed direction. On this chain the corrected audio is not an improvement.

Quality. Mean predicted overall quality across the segments went 2.45 → 2.85 (+0.40) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.2354normalised 0.453 → 0.990identity cos to seg 1 0.741 → 0.616 -0.125identity cos neighbours 0.741 → 0.616d_b rescored +0.546 → -0.043d_a rescored +0.546 → -0.043d_a mined 1.235d_b mined 0.537min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_02FqQjJyMtytotal 13.4schain gain +3.2 dBseam step 1.3 dBcrossfades 100 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an elderly masculine voice · neutral-toned, dark, slightly thin, below-average recording, quiet background, slow, relaxed, steady
(emotional numbness, contemplation · lethargic, narrow pitch range, whispered, monologue) So, ich würd mal sagen, Willkommen zu einer neuen Aufnahme, außerhalb des Streams.
full caption & clip details
An elderly masculine voice; delivery is lethargic, slow, relaxed, steady; timbre is neutral-toned, dark, smooth, slightly thin; slurred, frequent disfluency, narrow pitch range, audible breath; affect is mildly negative, submissive, neutral openness; reads as emotional numbness, contemplation; style: whispered, monologue; below-average recording, quiet background; genuineness 4.3/6; vocal-burst blend 0.0/10; 9.0s, DE.
DE_02FqQjJyMty_W000000 · in -18.4 dBFS · gain -1.6 dB · emolia-00215
(sadness, helplessness, distress · very low-energy, fairly narrow pitch, whispered, casual) In (low mumble) den Videos habe ich ja nicht mal mit dabei. Würde eigentlich Sinn machen.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is neutral-toned, dark, fairly smooth, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, submissive, neutral openness; reads as sadness, helplessness, distress; style: whispered, casual; below-average recording, quiet background; genuineness 5.7/6; vocal-burst blend 0.8/10; 4.6s, DE.
DE_02FqQjJyMty_W000002 · in -17.1 dBFS · gain -2.9 dB · emolia-00215
Distress rising ↑identity −0.00 emotion 11 %   sad-Distress-S3-k2 · #6

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.12. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.12.

On the corpus-wide percentile scale those become 0.44, 0.98 — a total move of +0.54.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 11 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.802 before conversion and 0.798 after — it fell by 0.004. Neighbour-to-neighbour the worst pair went 0.802 → 0.798. (The earlier render, with segment 1 left raw, scores 0.511 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.542 in the original and +0.061 after conversion — 11 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.46 → 2.82 (+0.37) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.1172normalised 0.453 → 0.986identity cos to seg 1 0.802 → 0.798 -0.004identity cos neighbours 0.802 → 0.798d_b rescored +0.542 → +0.061d_a rescored +0.542 → +0.061d_a mined 1.117d_b mined 0.533min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0PbfumbfjeQtotal 10.6schain gain +2.1 dBseam step 0.0 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, slightly relaxed, some disfluency, average clarity, moderate pitch range, light breath
(fatigue exhaustion · normal-paced, normally alert, moderately variable, casual) Karten und so weiter. Ich mach's gleich auf und zeig's euch richtig, aber ich zeig's einfach nur so kurz und dann.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion; style: casual, conversational; below-average recording, quiet background; genuineness 5.0/6; vocal-burst blend 2.8/10; 5.3s, DE.
DE_0PbfumbfjeQ_W000000 · in -19.7 dBFS · gain -0.3 dB · emolia-00161
(helplessness, sadness, distress · measured, subdued, fairly steady, monologue) Jetzt werde ich auf jeden Fall keine weiteren mehr bestellen. Das muss man jetzt erstmal alles verarbeiten, was ich hier.
full caption & clip details
An adult feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as helplessness, sadness, distress; style: monologue, ASMR; good recording, no background noise; genuineness 3.1/6; vocal-burst blend 0.7/10; 5.5s, DE.
DE_0PbfumbfjeQ_W000005 · in -18.4 dBFS · gain -1.6 dB · emolia-00161
Distress rising ↑identity −0.11 emotion 100 %   sad-Distress-S3-k2 · #7

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.36. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.36.

On the corpus-wide percentile scale those become 0.44, 0.99 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 17 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.954 before conversion and 0.839 after — it fell by 0.115. Neighbour-to-neighbour the worst pair went 0.954 → 0.839. (The earlier render, with segment 1 left raw, scores 0.872 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.550 in the original and +0.548 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.73 → 3.03 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.3613normalised 0.453 → 0.993identity cos to seg 1 0.954 → 0.839 -0.115identity cos neighbours 0.954 → 0.839d_b rescored +0.550 → +0.548d_a rescored +0.550 → +0.548d_a mined 1.361d_b mined 0.540min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0S219bZ30GQtotal 17.0schain gain +2.8 dBseam step 1.1 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(steady, no disfluency, newsreading, formal) Zusammen wollen wir uns verschiedene Begriffe aus dem queeren und feministischen Spektrum anschauen, Hintergründe zu den Themen beleuchten und Fragen klären.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.2/10; 7.3s, DE.
DE_0S219bZ30GQ_W000000 · in -22.0 dBFS · gain +2.0 dB · emolia-00097
(fear, sadness, pain · fairly steady, little disfluency, narration, formal) Es gibt natürlich auch Fälle, in denen Kinder operiert werden müssen, weil sie sonst starke Schmerzen hätten oder sterben würden. Dagegen haben die Verbände von intersexuellen Menschen auch nichts.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, sadness, pain; style: narration, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 9.9s, DE.
DE_0S219bZ30GQ_W000030 · in -22.7 dBFS · gain +2.7 dB · emolia-00097
Distress rising ↑identity +0.50 emotion 92 %   sad-Distress-S3-k2 · #8

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.00. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.00.

On the corpus-wide percentile scale those become 0.44, 0.98 — a total move of +0.54.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 30 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.303 before conversion and 0.802 after — it rose by 0.499. Neighbour-to-neighbour the worst pair went 0.303 → 0.802. (The earlier render, with segment 1 left raw, scores 0.801 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.536 in the original and +0.496 after conversion — 92 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.95 → 3.39 (+0.44) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.0039normalised 0.453 → 0.981identity cos to seg 1 0.303 → 0.802 +0.499identity cos neighbours 0.303 → 0.802d_b rescored +0.536 → +0.496d_a rescored +0.536 → +0.496d_a mined 1.004d_b mined 0.528min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0uIkaC-rBuItotal 29.6schain gain +2.2 dBseam step 0.2 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, frequent disfluency, average clarity
(contemplation, sourness, bitterness · normal-paced, normally alert, neutral tension, conversational) ist in dem Sinne ein besonderes Projekt gewesen, denn erstmal waren alle Mentorinnen sehr, sehr skeptisch, ob das überhaupt klappt. 嗯, (low mumble) denn, (low mumble) äh, keiner aus dem Team kannte sich so wirklich mit Mindstorms aus. Wir haben alle gesagt, so, ja, nee, äh, (ahem) macht's auch lieber so, nee, könntest auch so machen. Aber nein, sie sind stur geblieben, sind bei ihrer Idee geblieben, 嗯, (low mumble)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as contemplation, sourness, bitterness; style: conversational, casual; average recording, quiet background; genuineness 5.4/6; vocal-burst blend 5.5/10; 19.1s, DE.
DE_0uIkaC-rBuI_W000000 · in -17.3 dBFS · gain -2.7 dB · emolia-00254
(helplessness, distress, confusion · slow, very low-energy, relaxed, whispered) Man kann auch alleine spielen oder mit ganz vielen anderen Leuten. Man spielt nach seinen eigenen Regeln und baut Welten, die
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, submissive, neutral openness; reads as helplessness, distress, confusion; style: whispered, monologue; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.0/10; 10.7s, DE.
DE_0uIkaC-rBuI_W000005 · in -15.2 dBFS · gain -4.8 dB · emolia-00254
Distress rising ↑identity +0.11 emotion REVERSED   sad-Distress-S3-k2 · #9

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.08. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.08.

On the corpus-wide percentile scale those become 0.44, 0.98 — a total move of +0.54.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 20 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.608 before conversion and 0.714 after — it rose by 0.106. Neighbour-to-neighbour the worst pair went 0.608 → 0.714. (The earlier render, with segment 1 left raw, scores 0.716 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.540 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.47 → 2.85 (+0.38) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.0781normalised 0.453 → 0.984identity cos to seg 1 0.608 → 0.714 +0.106identity cos neighbours 0.608 → 0.714d_b rescored +0.540 → +0.000d_a rescored +0.540 → +0.000d_a mined 1.078d_b mined 0.532min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0ytNrKylOrctotal 19.7schain gain +2.4 dBseam step 1.2 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an elderly somewhat feminine voice · slightly warm, dark, smooth, slow, very low-energy, relaxed, steady, frequent disfluency
(sexual lust, contemplation, fatigue exhaustion · somewhat unclear, minimal breath, whispered, ASMR) Und so lasse dich tragen von den Energiewellen, welche sich bereits zu dir bewegen. Es sind die Energiewellen, welche über meinem Kanal zu dir fließen.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, smooth, balanced body; somewhat unclear, frequent disfluency, narrow pitch range, minimal breath; affect is mildly positive, submissive, neutral openness; reads as sexual lust, contemplation, fatigue exhaustion; style: whispered, ASMR; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 3.3/10; 15.1s, DE.
DE_0ytNrKylOrc_W000000 · in -16.7 dBFS · gain -3.3 dB · emolia-00215
(contemplation, longing, relief · slurred, audible breath, whispered, ASMR) Lasse dich tragen. Lasse dich heilen.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, smooth, very full; slurred, frequent disfluency, narrow pitch range, audible breath; affect is mildly positive, submissive, vulnerable; reads as contemplation, longing, relief; style: whispered, ASMR; very good recording, no background noise; genuineness 1.5/6; vocal-burst blend 6.0/10; 4.8s, DE.
DE_0ytNrKylOrc_W000001 · in -16.1 dBFS · gain -3.9 dB · emolia-00215
Distress rising ↑identity −0.11 emotion REVERSED   sad-Distress-S3-k2 · #10

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.11. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.11.

On the corpus-wide percentile scale those become 0.44, 0.98 — a total move of +0.54.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 21 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.938 before conversion and 0.827 after — it fell by 0.110. Neighbour-to-neighbour the worst pair went 0.938 → 0.827. (The earlier render, with segment 1 left raw, scores 0.752 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.541 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.85 → 3.11 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.1113normalised 0.453 → 0.986identity cos to seg 1 0.938 → 0.827 -0.110identity cos neighbours 0.938 → 0.827d_b rescored +0.541 → +0.000d_a rescored +0.541 → +0.000d_a mined 1.111d_b mined 0.533min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_12MSkDIQFvktotal 20.7schain gain -2.2 dBseam step 2.8 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an elderly somewhat feminine voice · neutral-toned, slightly dark, slightly rough, very low-energy, relaxed, fairly steady, frequent disfluency, somewhat unclear
(fatigue exhaustion, contemplation, longing · slow, monologue, ASMR) Und uns gewarnt, ja, nicht hinausschauen und ruhig bleiben und, (ahem) äh, und wir Kinder haben das nicht verstanden und,
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, submissive, neutral openness; reads as fatigue exhaustion, contemplation, longing; style: monologue, ASMR; below-average recording, some background noise; genuineness 3.0/6; vocal-burst blend 0.9/10; 9.5s, DE.
DE_12MSkDIQFvk_W000002 · in -16.8 dBFS · gain -3.2 dB · emolia-00003
(helplessness, fatigue exhaustion, distress · measured, ASMR, whispered) sind wir nicht ruhig geblieben, immer wieder bei die Fenster auße geschaut und beobachtet, was los ist. Bis wir sahen, dass sie uns relieferten,
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, submissive, neutral openness; reads as helplessness, fatigue exhaustion, distress; style: ASMR, whispered; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 0.7/10; 11.5s, DE.
DE_12MSkDIQFvk_W000003 · in -19.6 dBFS · gain -0.4 dB · emolia-00003
Distress rising ↑identity −0.07 emotion REVERSED   sad-Distress-S3-k2 · #11

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.02. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.02.

On the corpus-wide percentile scale those become 0.44, 0.98 — a total move of +0.54.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 12 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.907 before conversion and 0.841 after — it fell by 0.066. Neighbour-to-neighbour the worst pair went 0.907 → 0.841. (The earlier render, with segment 1 left raw, scores 0.811 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.537 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 3.03 → 3.15 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.0205normalised 0.453 → 0.982identity cos to seg 1 0.907 → 0.841 -0.066identity cos neighbours 0.907 → 0.841d_b rescored +0.537 → +0.000d_a rescored +0.537 → +0.000d_a mined 1.020d_b mined 0.529min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_163DYBR_hRgtotal 11.8schain gain +2.2 dBseam step 0.3 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(steady, formal, monologue) Startschuss für DSDS. Nicht mehr lange, und Florian Silbereisen-Fans kommen auch bei RTL auf ihre Kosten
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.9s, DE.
DE_163DYBR_hRg_W000000 · in -16.4 dBFS · gain -3.6 dB · emolia-00161
(confusion, helplessness, sadness · fairly steady, storytelling, formal) Ich sitze hier und weiß genauso wenig wie ihr, was heute passiert.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as confusion, helplessness, sadness; style: storytelling, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.5/10; 4.1s, DE.
DE_163DYBR_hRg_W000011 · in -16.4 dBFS · gain -3.6 dB · emolia-00161
Distress rising ↑identity −0.00 emotion 102 %   sad-Distress-S3-k2 · #12

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.15. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.15.

On the corpus-wide percentile scale those become 0.44, 0.99 — a total move of +0.54.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 13 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.793 before conversion and 0.793 after — it fell by 0.000. Neighbour-to-neighbour the worst pair went 0.793 → 0.793. (The earlier render, with segment 1 left raw, scores 0.683 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.543 in the original and +0.554 after conversion — 102 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.85 → 3.00 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.1533normalised 0.453 → 0.987identity cos to seg 1 0.793 → 0.793 -0.000identity cos neighbours 0.793 → 0.793d_b rescored +0.543 → +0.554d_a rescored +0.543 → +0.554d_a mined 1.153d_b mined 0.534min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1A7Boronwiktotal 13.1schain gain +1.2 dBseam step 0.1 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an elderly somewhat feminine voice · slightly warm, smooth, balanced body, no background noise, slow, very low-energy, relaxed, steady
(contentment, affection, relief · no disfluency, clear, narrow pitch range, ASMR) Es ist mir ein Bedürfnis, heute mit dir zusammen, mit euch allen zusammen.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, neutral-bright, smooth, balanced body; clear, no disfluency, narrow pitch range, audible breath; affect is mildly positive, submissive, neutral openness; reads as contentment, affection, relief; style: ASMR, monologue; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.4/10; 8.7s, DE.
DE_1A7Boronwik_W000000 · in -18.1 dBFS · gain -1.9 dB · emolia-00254
(shame, sadness, relief · frequent disfluency, slurred, fairly narrow pitch, whispered) Und ich bin mir meiner Verantwortung bewusst.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, smooth, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, submissive, vulnerable; reads as shame, sadness, relief; style: whispered, monologue; very good recording, no background noise; genuineness 1.6/6; vocal-burst blend 3.8/10; 4.6s, DE.
DE_1A7Boronwik_W000006 · in -19.8 dBFS · gain -0.2 dB · emolia-00254
Distress rising ↑identity +0.37 emotion REVERSED   sad-Distress-S3-k2 · #13

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.24. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.24.

On the corpus-wide percentile scale those become 0.44, 0.99 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 13 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.166 before conversion and 0.533 after — it rose by 0.366. Neighbour-to-neighbour the worst pair went 0.166 → 0.533. (The earlier render, with segment 1 left raw, scores 0.453 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.546 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.70 → 2.84 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.2402normalised 0.453 → 0.990identity cos to seg 1 0.166 → 0.533 +0.366identity cos neighbours 0.166 → 0.533d_b rescored +0.546 → +0.000d_a rescored +0.546 → +0.000d_a mined 1.240d_b mined 0.537min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1Pd71w8hAhItotal 12.6schain gain +1.3 dBseam step 3.6 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, quiet background, neutral tension, moderately variable, somewhat unclear, normal breath
(embarrassment, fatigue exhaustion, infatuation · normal-paced, normally alert, some disfluency, casual) Ich drück mich immer falsch auf, aus, in dieser, also, filterraum, ich muss auch mal die richtige Sprache lernen.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, normal breath; affect is positive, slightly submissive, slightly vulnerable; reads as embarrassment, fatigue exhaustion, infatuation; style: casual, conversational; below-average recording, quiet background; genuineness 5.3/6; vocal-burst blend 4.9/10; 6.3s, DE.
DE_1Pd71w8hAhI_W000000 · in -21.4 dBFS · gain +1.4 dB · emolia-00077
(helplessness, fatigue exhaustion, distress · measured, very low-energy, frequent disfluency, conversational) (low mumble) euhm, also, ich hab so das Gefühl, ich lass mir da auch Zeit, ja, das drängt mich niemand. (breathy giggle)
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, vulnerable; reads as helplessness, fatigue exhaustion, distress; style: conversational, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 1.9/10; 6.5s, DE.
DE_1Pd71w8hAhI_W000014 · in -17.7 dBFS · gain -2.3 dB · emolia-00077
Distress rising ↑identity +0.35 emotion 98 %   sad-Distress-S3-k2 · #14

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.02. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.02.

On the corpus-wide percentile scale those become 0.44, 0.98 — a total move of +0.54.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 25 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.081 before conversion and 0.434 after — it rose by 0.353. Neighbour-to-neighbour the worst pair went 0.081 → 0.434. (The earlier render, with segment 1 left raw, scores 0.349 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.537 in the original and +0.528 after conversion — 98 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.76 → 3.01 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.0205normalised 0.453 → 0.982identity cos to seg 1 0.081 → 0.434 +0.353identity cos neighbours 0.081 → 0.434d_b rescored +0.537 → +0.528d_a rescored +0.537 → +0.528d_a mined 1.020d_b mined 0.529min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1ZMIiMMqy9utotal 24.5schain gain +2.4 dBseam step 2.2 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(steady, no disfluency, fairly narrow pitch, authoritative) Antrag der Fraktion Bündnis 90, die GRÜNEN.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, no audible breath; affect is neutral, dominant, slightly guarded; no dominant emotion; style: authoritative, formal; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.6/10; 3.1s, DE.
DE_1ZMIiMMqy9u_W000000 · in -20.9 dBFS · gain +0.9 dB · emolia-00123
(disappointment, bitterness, sadness · fairly steady, almost no disfluency, moderate pitch range, newsreading) Wo rechtsextreme Gewalttaten die Gesundheit und das Leben von Migrantinnen und Migranten, von Antifaschistinnen und Antifaschisten bedrohen. 600 gemeldete Straftaten, die in den Bereich politisch motivierte Kriminalität rechts zuzuordnen sind, gab es allein im Jahr 2017, darunter 17 Gewalttaten.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as disappointment, bitterness, sadness; style: newsreading, formal; average recording, quiet background; genuineness 0.5/6; vocal-burst blend 0.7/10; 21.6s, DE.
DE_1ZMIiMMqy9u_W000072 · in -21.8 dBFS · gain +1.8 dB · emolia-00123
Distress rising ↑identity −0.06 emotion 97 %   sad-Distress-S3-k2 · #15

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.33. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.33.

On the corpus-wide percentile scale those become 0.44, 0.99 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 27 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.875 before conversion and 0.814 after — it fell by 0.061. Neighbour-to-neighbour the worst pair went 0.875 → 0.814. (The earlier render, with segment 1 left raw, scores 0.754 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.549 in the original and +0.534 after conversion — 97 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 3.14 → 3.33 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.3271normalised 0.453 → 0.992identity cos to seg 1 0.875 → 0.814 -0.061identity cos neighbours 0.875 → 0.814d_b rescored +0.549 → +0.534d_a rescored +0.549 → +0.534d_a mined 1.327d_b mined 0.539min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1aPh9BCqtRutotal 26.3schain gain +2.2 dBseam step 3.2 dBcrossfades 100 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, energised, neutral tension
(disgust, interest, teasing · frequent disfluency, very clear, very wide pitch range, storytelling) 3G Symmension 3. Ihr wisst Bescheid. Ganz genau. Und heute befinden wir uns, wenn ich das hier alles richtig sehe, und mich gut informiert hab, in Pyra Media. Ja, (ahem) ganz genau.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; very clear, frequent disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as disgust, interest, teasing; style: storytelling, playful; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 1.9/10; 14.1s, DE.
DE_1aPh9BCqtRu_W000000 · in -15.4 dBFS · gain -4.6 dB · emolia-00022
(confusion, helplessness, doubt · some disfluency, average clarity, wide pitch range, casual) Was ist das denn? Auf einmal spot eine Kugel und ich kann nicht weiter Staubsaugermeister machen hier. Schön, dass es dafür ein Kristall gibt. Ich weiß nicht, wo ich laufen darf.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is elated, slightly dominant, neutral openness; reads as confusion, helplessness, doubt; style: casual, dramatic; average recording, quiet background; mildly explicit content; genuineness 3.8/6; vocal-burst blend 0.0/10; 12.3s, DE.
DE_1aPh9BCqtRu_W000028 · in -16.8 dBFS · gain -3.2 dB · emolia-00022
Distress rising ↑identity +0.38 emotion 94 %   sad-Distress-S3-k2 · #16

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.06. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.06.

On the corpus-wide percentile scale those become 0.44, 0.98 — a total move of +0.54.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 24 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.448 before conversion and 0.827 after — it rose by 0.379. Neighbour-to-neighbour the worst pair went 0.448 → 0.827. (The earlier render, with segment 1 left raw, scores 0.659 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.539 in the original and +0.506 after conversion — 94 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.86 → 3.16 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.0566normalised 0.453 → 0.984identity cos to seg 1 0.448 → 0.827 +0.379identity cos neighbours 0.448 → 0.827d_b rescored +0.539 → +0.506d_a rescored +0.539 → +0.506d_a mined 1.057d_b mined 0.531min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1c4KtO2dsKEtotal 23.2schain gain +2.9 dBseam step 0.3 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, quiet background, normal-paced, normally alert
(teasing, amusement, intoxication altered states of consciousness · neutral tension, moderately variable, wide pitch range, conversational) Es, es, es sind auch, (ahem) äh, andere Sachen gemeint. Das war jetzt bloß ein Beispiel. Aber, (childlike giggle) äh, also ich glaube, sexy Klamotten, meinen Klamottenstil nicht nennen. Einfach nur, oh, ist bequem.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as teasing, amusement, intoxication altered states of consciousness; style: conversational, playful; good recording, quiet background; genuineness 4.5/6; vocal-burst blend 0.9/10; 13.5s, DE.
DE_1c4KtO2dsKE_W000000 · in -12.4 dBFS · gain -7.5 dB · emolia-00167
(shame, sadness, helplessness · slightly relaxed, fairly steady, moderate pitch range, conversational) Ich hab schon mal wissentlich geflirtet, aber mir ist es auch schon mal passiert, dass irgendwie ein Verhalten von mir als flirten gedeutet wurde und von mir da nicht so gemeint war. Wir haben hier noch eine letzte.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as shame, sadness, helplessness; style: conversational, casual; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.0/10; 9.9s, DE.
DE_1c4KtO2dsKE_W000024 · in -13.0 dBFS · gain -7.0 dB · emolia-00167
Distress rising ↑identity +0.51 emotion 90 %   sad-Distress-S3-k2 · #17

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.85. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.85.

On the corpus-wide percentile scale those become 0.44, 1.00 — a total move of +0.56.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 12 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.267 before conversion and 0.774 after — it rose by 0.507. Neighbour-to-neighbour the worst pair went 0.267 → 0.774. (The earlier render, with segment 1 left raw, scores 0.715 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.556 in the original and +0.501 after conversion — 90 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.88 → 3.08 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.8447normalised 0.453 → 0.998identity cos to seg 1 0.267 → 0.774 +0.507identity cos neighbours 0.267 → 0.774d_b rescored +0.556 → +0.501d_a rescored +0.556 → +0.501d_a mined 1.845d_b mined 0.545min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1eb_VRL8MFItotal 11.5schain gain +2.3 dBseam step 0.8 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normal-paced, normally alert, fairly steady
(affection · slightly relaxed, little disfluency, clear, authoritative) Tage der Ermutigung. Ich freue mich, dass du wieder eingeschaltet hast, dass du wieder dabei bist.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as affection; style: authoritative, storytelling; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.0/10; 4.9s, DE.
DE_1eb_VRL8MFI_W000000 · in -19.3 dBFS · gain -0.7 dB · emolia-00216
(pain, relief, helplessness · neutral tension, some disfluency, average clarity, conversational) Ich wurde so krank, dass ich nicht mehr laufen konnte, dass ich mich nicht mehr bewegen konnte und ich große Schmerzen hatte im ganzen Körper.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as pain, relief, helplessness; style: conversational, casual; good recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.3/10; 6.9s, DE.
DE_1eb_VRL8MFI_W000004 · in -20.9 dBFS · gain +0.9 dB · emolia-00216
Distress rising ↑identity +0.17 emotion REVERSED   sad-Distress-S3-k2 · #18

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.36. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.36.

On the corpus-wide percentile scale those become 0.44, 0.99 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 11 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.586 before conversion and 0.755 after — it rose by 0.170. Neighbour-to-neighbour the worst pair went 0.586 → 0.755. (The earlier render, with segment 1 left raw, scores 0.546 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.550 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.69 → 2.95 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.3594normalised 0.453 → 0.993identity cos to seg 1 0.586 → 0.755 +0.170identity cos neighbours 0.586 → 0.755d_b rescored +0.550 → +0.000d_a rescored +0.550 → +0.000d_a mined 1.359d_b mined 0.540min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1g2ByepregYtotal 10.4schain gain +1.9 dBseam step 2.3 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · thin, below-average recording, quiet background, fast, some disfluency
(affection · normally alert, fully relaxed, fairly steady, casual) Rock, natürlich, gib ihm die Frisur. Das war kleiner Jojo.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, fully relaxed, fairly steady; timbre is neutral-toned, dark, slightly rough, thin; slurred, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as affection; style: casual, conversational; below-average recording, quiet background; genuineness 5.9/6; vocal-burst blend 1.9/10; 3.5s, DE.
DE_1g2ByepregY_W000000 · in -15.3 dBFS · gain -4.7 dB · emolia-00161
(fear, distress, astonishment surprise · highly aroused, tense, moderately variable, casual) Ja. What the fuck? Das Ding ist halt so, du spritzt das Eimer ab und es ist kaputt. So ein Zweig gab es in irgendeinem Modus.
full caption & clip details
A young adult somewhat masculine voice; delivery is highly aroused, fast, tense, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, very wide pitch range, normal breath; affect is elated, very dominant, neutral openness; reads as fear, distress, astonishment surprise; style: casual, dramatic; below-average recording, quiet background; genuineness 4.3/6; vocal-burst blend 1.7/10; 7.1s, DE.
DE_1g2ByepregY_W000019 · in -18.5 dBFS · gain -1.5 dB · emolia-00161
Distress rising ↑identity +0.06 emotion 100 %   sad-Distress-S3-k2 · #19

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.50. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.50.

On the corpus-wide percentile scale those become 0.44, 0.99 — a total move of +0.55.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 36 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.627 before conversion and 0.689 after — it rose by 0.062. Neighbour-to-neighbour the worst pair went 0.627 → 0.689. (The earlier render, with segment 1 left raw, scores 0.658 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.553 in the original and +0.552 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.85 → 3.18 (+0.32) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.5normalised 0.453 → 0.995identity cos to seg 1 0.627 → 0.689 +0.062identity cos neighbours 0.627 → 0.689d_b rescored +0.553 → +0.552d_a rescored +0.553 → +0.552d_a mined 1.500d_b mined 0.542min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1gZLllrev9ctotal 35.8schain gain +3.3 dBseam step 1.0 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, measured, slightly relaxed, average clarity
(normally alert, fairly steady, frequent disfluency, didactic) Wie bereits in einem vergangenen Video erwähnt, ist die Frage nach gut oder schlecht schwer zu beantworten. Diese Frage ist aber sehr wichtig.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.0/10; 8.4s, DE.
DE_1gZLllrev9c_W000000 · in -19.9 dBFS · gain -0.1 dB · emolia-00259
(bitterness, helplessness, sourness · subdued, steady, some disfluency, monologue) Dies lässt sich als evolutionärer Prozess betrachten, in dem die Moral überlebt, welche ihren Trägern die besten Voraussetzungen zum Überleben gibt und jede Moral durch beispielsweise Missverständnisse langsam mutiert. Da die Existenz ihrer Träger das Axiom meiner Moral ist, ist sie zwingend immer am besten dafür geeignet und wird sich, ob nun durch mich oder nicht, irgendwann durchsetzen.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, fairly guarded; reads as bitterness, helplessness, sourness; style: monologue, narration; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 0.1/10; 27.7s, DE.
DE_1gZLllrev9c_W000017 · in -19.2 dBFS · gain -0.8 dB · emolia-00259
Distress rising ↑identity +0.06 emotion 98 %   sad-Distress-S3-k2 · #20

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Distress.

The raw scorer output across the chain is 0.00, 1.09. In the first clip the scorer found no Distress whatsoever (0.00); by the last it is at 1.09.

On the corpus-wide percentile scale those become 0.44, 0.98 — a total move of +0.54.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Distress, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Distress sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 27 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.828 before conversion and 0.887 after — it rose by 0.059. Neighbour-to-neighbour the worst pair went 0.828 → 0.887. (The earlier render, with segment 1 left raw, scores 0.739 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.540 in the original and +0.531 after conversion — 98 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.65 → 3.17 (+0.52) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores 0.0 → 1.0859normalised 0.453 → 0.985identity cos to seg 1 0.828 → 0.887 +0.059identity cos neighbours 0.828 → 0.887d_b rescored +0.540 → +0.531d_a rescored +0.540 → +0.531d_a mined 1.086d_b mined 0.532min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_2R_mAsRw2RItotal 26.5schain gain -0.6 dBseam step 2.8 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, fairly smooth, balanced body, good recording, normally alert, slightly relaxed, some disfluency, average clarity
(elation, pleasure ecstasy, contentment · brisk, moderately variable, wide pitch range, playful) Hi, hallo, ich bin's Alena. Ich freue mich, dass ich dir heute meine Morgenroutine zeigen darf. Diese Morgenroutine ist jeden Montag bis jeden Freitag, denn ich gehe arbeiten. Mein Tag beginnt um 5.15 Uhr oder auch 5.30 Uhr, je nachdem, ob ich mir noch wie heute die Haare machen muss.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as elation, pleasure ecstasy, contentment; style: playful, conversational; good recording, quiet background; genuineness 1.8/6; vocal-burst blend 1.4/10; 16.1s, DE.
DE_2R_mAsRw2RI_W000002 · in -13.8 dBFS · gain -6.2 dB · emolia-00258
(fatigue exhaustion, longing, contentment · normal-paced, fairly steady, moderate pitch range, casual) Dehne ich mich am Morgen immer ganz gerne und entspanne nochmal so ein bisschen. Dehne auch meinen Schultern und Nackenbereich dadurch, dass ich ja einen ganzen Tag am Schreibtisch am Computer sitze auf der Arbeit.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as fatigue exhaustion, longing, contentment; style: casual, monologue; good recording, no background noise; genuineness 3.5/6; vocal-burst blend 0.7/10; 10.7s, DE.
DE_2R_mAsRw2RI_W000033 · in -13.6 dBFS · gain -6.4 dB · emolia-00258