rescue rule sad-Helplessness-BASE-k3 — voice-corrected

Helplessness under rescue rule BASE, k=3. CONTROL: Helplessness has a gap of only 0.105, below the 0.25 step cap, so the STRICT rule already works here -- it was never blocked. 88.8 % of clips score at or below zero on this emotion and the largest gap on its normalised axis is 0.105 (narrower than the 0.25 step cap). This rule found 19,693 chains over 40,000 tracks; the strict rule found 19,693 at k=3.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_sad-Helplessness-BASE-k3.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
These are not strict-rule trajectories. They come from a deliberately looser rule, built to recover examples on an emotion the strict rule cannot reach. What rule BASE changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25. What it costs: Nothing -- this is the strict rule everything else is measured against. Full explanation →
20chains converted
40segments re-voiced
0.753 → 0.797median worst-to-anchor identity cosine
96 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Helplessness rising ↑identity +0.01 emotion REVERSED   sad-Helplessness-BASE-k3 · #1

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.41, 0.41, 0.76 — a total move of +0.35.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.35 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 32 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.663 before conversion and 0.677 after — it rose by 0.014. Neighbour-to-neighbour the worst pair went 0.663 → 0.664. (The earlier render, with segment 1 left raw, scores 0.502 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Helplessness moved +0.353 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.88 → 3.15 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0044 → -0.0042 → -0.0039normalised 0.391 → 0.598 → 0.831identity cos to seg 1 0.663 → 0.677 +0.014identity cos neighbours 0.663 → 0.664d_b rescored +0.353 → +0.000d_a rescored +0.353 → +0.000d_a mined 0.001d_b mined 0.439min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_--z5fTsHDactotal 30.9schain gain +0.3 dBseam step 0.9 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, moderate pitch range, light breath
(awe, contentment · measured, steady, some disfluency, didactic) Und dabei sind wir zu der Einschätzung gekommen, dass der Herr Dr. Markus Bühler wahrscheinlich eine ödipale Fixierung auf seine Mutter hat. Vermutlich hat seine Mutter ihn so erzogen, dass er seinen Vater
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, contentment; style: didactic, monologue; average recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.3/10; 16.1s, DE.
DE_--z5fTsHDac_W000009 · in -18.6 dBFS · gain -1.4 dB · emolia-00037
(normal-paced, fairly steady, no disfluency, formal) Es ist natürlich schwierig, ja, hat auch der Psychologe gesagt, mit dem geredet hab.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 0.0/10; 4.1s, DE.
DE_--z5fTsHDac_W000017 · in -15.1 dBFS · gain -4.9 dB · emolia-00037
(concentration, fatigue exhaustion · measured, fairly steady, some disfluency, monologue) Und natürlich könnt ihr auch Spenden für Freifarmen, um unsere Sache voranzubringen. Den Link zu den Spendeninformationen findet ihr jetzt im Abspann. Danke nochmal fürs Zuschauen und tschüss zusammen.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, fatigue exhaustion; style: monologue, casual; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.6/10; 11.1s, DE.
DE_--z5fTsHDac_W000020 · in -19.3 dBFS · gain -0.7 dB · emolia-00037
Helplessness rising ↑identity +0.01 emotion —   sad-Helplessness-BASE-k3 · #2

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.41, 0.41, 0.41 — a total move of +0.00.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.00 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 58 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.911 before conversion and 0.917 after — it rose by 0.006. Neighbour-to-neighbour the worst pair went 0.912 → 0.917. (The earlier render, with segment 1 left raw, scores 0.843 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The re-scored move on Helplessness is +0.000 before and +0.000 after, but the before-value is too close to zero for a retention ratio to mean anything on this chain.

Quality. Mean predicted overall quality across the segments went 3.13 → 3.24 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0045 → -0.0044 → -0.0042normalised 0.289 → 0.391 → 0.598identity cos to seg 1 0.911 → 0.917 +0.006identity cos neighbours 0.912 → 0.917d_b rescored +0.000 → +0.000d_a rescored +0.000 → +0.000d_a mined 0.000d_b mined 0.310min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-6F1Ss25DJQtotal 57.0schain gain +3.3 dBseam step 1.0 dBcrossfades 100/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(jealousy and envy · monologue, didactic) Ergibt also 173 Dollar, die jetzt hier quasi aus dem Crowdloan rausgekommen sind und das invest. Das lag zum Zeitpunkt der (ahem) Crowdloan Teilnahme bei 530 Dollar. Also haben wir hier ungefähr 32, 33 Prozent, ein Drittel.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy; style: monologue, didactic; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.1/10; 20.3s, DE.
DE_-6F1Ss25DJQ_W000026 · in -19.3 dBFS · gain -0.7 dB · emolia-00037
(sourness, teasing, concentration · monologue, casual) Außer Solarflare findet irgendeine gute Anwendung. Hier gibt es zum Beispiel das Staking. Da kann man also seine Flares westen für einige Zeit. Und, (ahem) äh, wenn man das macht, dann, (low mumble) ähm, bekommt man hier eine 840% LPA, die wahrscheinlich auch in Zukunft sinken wird. (low mumble) Und da, (low mumble) ähm,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, teasing, concentration; style: monologue, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 0.0/10; 17.3s, DE.
DE_-6F1Ss25DJQ_W000120 · in -17.9 dBFS · gain -2.1 dB · emolia-00037
(sourness, hope enthusiasm optimism, intoxication altered states of consciousness · casual, monologue) (ahem) äh, zum Beispiel auch hier Stablecoins, die, äh, (low mumble) haben wir, können wir, müssen wir irgendwie hin übertragen. Und das ist natürlich ein Thema für ein eigenes Video, weil das Ganze jetzt so lang wird. Und das werde ich jetzt hoffentlich schaffen, in den nächsten Tagen auch zu produzieren. Bis dahin wünsche ich euch erstmal viel Spaß mit den Glimmer-Tokens, viel Erfolg, macht's gut, tschüss!
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as sourness, hope enthusiasm optimism, intoxication altered states of consciousness; style: casual, monologue; good recording, quiet background; genuineness 3.5/6; vocal-burst blend 2.1/10; 19.7s, DE.
DE_-6F1Ss25DJQ_W000125 · in -19.1 dBFS · gain -0.9 dB · emolia-00037
Helplessness rising ↑identity +0.33 emotion 112 %   sad-Helplessness-BASE-k3 · #3

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, 0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at 0.00.

On the corpus-wide percentile scale those become 0.41, 0.76, 0.86 — a total move of +0.46.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Helplessness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Helplessness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 39 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.014 before conversion and 0.347 after — it rose by 0.333. Neighbour-to-neighbour the worst pair went 0.361 → 0.512. (The earlier render, with segment 1 left raw, scores 0.283 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Helplessness moved +0.456 in the original and +0.510 after conversion — 112 % of the delta retained, i.e. the move came out slightly larger after conversion than before.

Quality. Mean predicted overall quality across the segments went 2.85 → 3.06 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0042 → -0.0039 → 0.0041normalised 0.598 → 0.831 → 0.896identity cos to seg 1 0.014 → 0.347 +0.333identity cos neighbours 0.361 → 0.512d_b rescored +0.456 → +0.510d_a rescored +0.456 → +0.510d_a mined 0.008d_b mined 0.297min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-8Xv740j-8Qtotal 38.1schain gain +0.2 dBseam step 1.6 dBcrossfades 100/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady
(confusion, doubt · some disfluency, storytelling, casual) Ich würde sagen, da ist auf jeden Fall Bedarf der Verbesserung.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as confusion, doubt; style: storytelling, casual; good recording, no background noise; genuineness 2.4/6; vocal-burst blend 0.0/10; 3.7s, DE.
DE_-8Xv740j-8Q_W000019 · in -16.3 dBFS · gain -3.7 dB · emolia-00037
(affection · little disfluency, conversational, authoritative) Ja, dann bedanke ich mich ganz herzlich bei Ihnen, dass Sie heute hier sind und dass Sie auch zur Verfügung stehen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as affection; style: conversational, authoritative; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 0.0/10; 4.7s, DE.
DE_-8Xv740j-8Q_W000046 · in -15.3 dBFS · gain -4.7 dB · emolia-00037
(bitterness, jealousy and envy, pain · some disfluency, monologue, authoritative) Die erste Kundin ist bereits in Beratung. Die ersten Gespräche sind angebracht und damit wird auch der Webauftritt noch weiter dann optimiert. Und wir wünschen jetzt dir auch viel Erfolg dabei, wenn du dich in der Reinigungsbranche selbstständig machen möchtest oder jemand kennst. Du kannst gerne den Link unterhalb des Videos nutzen für ein kostenfreies Erstgespräch. Wir würden dann gemeinsam als Team zunächst sicherstellen, dass der AVGS dafür beantragt wird. Was ist ein AVGS?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as bitterness, jealousy and envy, pain; style: monologue, authoritative; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 1.3/10; 30.0s, DE.
DE_-8Xv740j-8Q_W000047 · in -19.1 dBFS · gain -0.9 dB · emolia-00037
Helplessness rising ↑identity −0.08 emotion —   sad-Helplessness-BASE-k3 · #4

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.41, 0.41, 0.41 — a total move of +0.00.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.00 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 30 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.867 before conversion and 0.786 after — it fell by 0.081. Neighbour-to-neighbour the worst pair went 0.815 → 0.682. (The earlier render, with segment 1 left raw, scores 0.705 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The re-scored move on Helplessness is +0.000 before and +0.437 after, but the before-value is too close to zero for a retention ratio to mean anything on this chain.

Quality. Mean predicted overall quality across the segments went 3.03 → 3.24 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0044 → -0.0042 → -0.0041normalised 0.391 → 0.598 → 0.691identity cos to seg 1 0.867 → 0.786 -0.081identity cos neighbours 0.815 → 0.682d_b rescored +0.000 → +0.437d_a rescored +0.000 → +0.437d_a mined 0.000d_b mined 0.300min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-8c1BTgFP3ktotal 29.1schain gain +3.0 dBseam step 0.6 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, clear, light breath
(contempt · measured, steady, almost no disfluency, didactic) Corona darf unsere Demokratie nicht endgültig zur Zerreißprobe bringen. Gleichzeitig gibt es neue Bedrohungen aus Russland. Osteuropa droht gar im Krieg.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt; style: didactic, monologue; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.0/10; 10.1s, DE.
DE_-8c1BTgFP3k_W000011 · in -21.5 dBFS · gain +1.5 dB · emolia-00247
(relief, emotional numbness · normal-paced, fairly steady, no disfluency, formal) mit Entschlossenheit reagieren, sagte Steinmeier dazu. Das hat mir sehr gut gefallen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, emotional numbness; style: formal, didactic; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 0.0/10; 5.0s, DE.
DE_-8c1BTgFP3k_W000012 · in -21.1 dBFS · gain +1.1 dB · emolia-00247
(sourness, disappointment, relief · normal-paced, fairly steady, little disfluency, monologue) Hoffen wir also, dass die Ampelregierung in Berlin nun endlich trittfast und am Ende nicht wieder der Bundespräsident und ehemalige Außenminister Steinmeier die Kastanen aus dem Feuer holen muss. Bitte bleibt gesund und Gott schütze euch.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, disappointment, relief; style: monologue, didactic; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.0/10; 14.4s, DE.
DE_-8c1BTgFP3k_W000014 · in -21.3 dBFS · gain +1.3 dB · emolia-00247
Helplessness rising ↑identity −0.00 emotion —   sad-Helplessness-BASE-k3 · #5

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.41, 0.41, 0.41 — a total move of +0.00.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.00 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 32 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.846 before conversion and 0.842 after — it fell by 0.004. Neighbour-to-neighbour the worst pair went 0.821 → 0.847. (The earlier render, with segment 1 left raw, scores 0.785 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The re-scored move on Helplessness is +0.000 before and +0.000 after, but the before-value is too close to zero for a retention ratio to mean anything on this chain.

Quality. Mean predicted overall quality across the segments went 2.94 → 3.20 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0045 → -0.0044 → -0.0042normalised 0.289 → 0.391 → 0.598identity cos to seg 1 0.846 → 0.842 -0.004identity cos neighbours 0.821 → 0.847d_b rescored +0.000 → +0.000d_a rescored +0.000 → +0.000d_a mined 0.000d_b mined 0.310min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-HcL6XTJTVgtotal 31.4schain gain +1.6 dBseam step 0.4 dBcrossfades 100/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, brisk, some disfluency, average clarity
(confusion, interest, jealousy and envy · normally alert, slightly relaxed, fairly steady, casual) Solltet ihr auch machen, damit ihr seht, was passiert. Ich klicke hier unten auf Create, dann auf Hinge und dann mittig auf diesen Block und mittig auf die Sphere. Damit haben wir eine Hinge erstellt.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as confusion, interest, jealousy and envy; style: casual, dramatic; good recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.7/10; 9.4s, DE.
DE_-HcL6XTJTVg_W000003 · in -16.0 dBFS · gain -4.0 dB · emolia-00047
(amusement, intoxication altered states of consciousness, contempt · energised, neutral tension, moderately variable, casual) And debounce is like false. Das hier ist einfach so simpler debounce, da er dafür sorgt, dass nicht permanent das Auto neu gespawnt wird, sondern nur alle fünf Sekunden gespawnt werden kann. Und hier drunter sagen wir jetzt einfach local neu ist gleich Auto Doppelpunkt clone neu.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as amusement, intoxication altered states of consciousness, contempt; style: casual, dramatic; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.0/10; 16.2s, DE.
DE_-HcL6XTJTVg_W000032 · in -16.1 dBFS · gain -3.9 dB · emolia-00047
(impatience and irritability, anger · energised, neutral tension, moderately variable, dramatic) können damit rumfahren. Es funktioniert auf jeden Fall. Wunderbar. Ich spring wieder raus aus dem Auto. Und wenn wir zum Spawn-Part wieder hinlaufen.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, anger; style: dramatic, playful; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.9/10; 6.1s, DE.
DE_-HcL6XTJTVg_W000036 · in -15.2 dBFS · gain -4.8 dB · emolia-00047
Helplessness rising ↑identity +0.02 emotion 93 %   sad-Helplessness-BASE-k3 · #6

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, 0.11. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at 0.11.

On the corpus-wide percentile scale those become 0.41, 0.41, 0.91 — a total move of +0.50.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Helplessness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Helplessness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 45 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.882 before conversion and 0.900 after — it rose by 0.018. Neighbour-to-neighbour the worst pair went 0.882 → 0.900. (The earlier render, with segment 1 left raw, scores 0.845 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Helplessness moved +0.496 in the original and +0.459 after conversion — 93 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.90 → 3.17 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0042 → -0.0041 → 0.1084normalised 0.598 → 0.691 → 0.927identity cos to seg 1 0.882 → 0.900 +0.018identity cos neighbours 0.882 → 0.900d_b rescored +0.496 → +0.459d_a rescored +0.496 → +0.459d_a mined 0.113d_b mined 0.328min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-NcjdeCDcbYtotal 43.9schain gain +3.5 dBseam step 0.5 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normal-paced, normally alert, slightly relaxed
(anger, malevolence malice · almost no disfluency, clear, didactic, formal) Das vierstöckige Gebäude wurde unterhalb der angrenzenden, stark befahrenen Hauptstraße gebaut und lag sehr zentral, gut zu erreichen, zu Fuß mit dem Auto, der Straßenbahn oder der angrenzenden Zugverbindung.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger, malevolence malice; style: didactic, formal; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.0/10; 11.7s, DE.
DE_-NcjdeCDcbY_W000004 · in -20.7 dBFS · gain +0.7 dB · emolia-00182
(concentration, jealousy and envy, shame · some disfluency, average clarity, monologue, formal) Man hatte zwar versucht, das Gebäude zu sanieren, allerdings in nur zwei Räumen. Warum der Rest auf der Strecke blieb, ist uns nicht bekannt. Im Inneren des Gebäudes fanden wir allerhand Unrat, Müll und diverse Hinterlassenschaften von Jugendlichen und Vandalen. Ich las auch etwas von einem Ballsaal, allerdings konnten wir einen solchen Raum weder zuordnen, noch überhaupt finden.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, jealousy and envy, shame; style: monologue, formal; good recording, quiet background; genuineness 1.8/6; vocal-burst blend 1.7/10; 20.4s, DE.
DE_-NcjdeCDcbY_W000006 · in -20.7 dBFS · gain +0.7 dB · emolia-00182
(disgust, fear, sourness · almost no disfluency, clear, formal, monologue) im ehemaligen Eingangsbereich zur Straßenseite hatte es vor einigen Jahren stark gebrannt weswegen es aufgrund der Witterung zu einem Dach- und Deckendurchbruch kam somit konnte man gerade den vorderen Teil gar nicht mehr betreten
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, fear, sourness; style: formal, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.1/10; 12.2s, DE.
DE_-NcjdeCDcbY_W000007 · in -21.3 dBFS · gain +1.3 dB · emolia-00182
Helplessness rising ↑identity +0.25 emotion REVERSED   sad-Helplessness-BASE-k3 · #7

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.41, 0.41, 0.76 — a total move of +0.35.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.35 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 24 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.387 before conversion and 0.640 after — it rose by 0.253. Neighbour-to-neighbour the worst pair went 0.387 → 0.640. (The earlier render, with segment 1 left raw, scores 0.574 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

The emotional move did not survive. Re-scored end to end, Helplessness moved +0.353 in the original and -0.353 after conversion — it changed direction. On this chain the corrected audio is not an improvement.

Quality. Mean predicted overall quality across the segments went 2.81 → 3.11 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0044 → -0.0042 → -0.004normalised 0.391 → 0.598 → 0.770identity cos to seg 1 0.387 → 0.640 +0.253identity cos neighbours 0.387 → 0.640d_b rescored +0.353 → -0.353d_a rescored +0.353 → -0.353d_a mined 0.000d_b mined 0.379min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-OLjkEFs1wktotal 22.8schain gain +0.3 dBseam step 1.1 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, normal-paced, normally alert, slightly relaxed
(sourness · fairly steady, some disfluency, average clarity, monologue) Oblikatorische Nullrunde führen, also nicht, dass man wieder versucht, solche Argumente durch die Hintertür. Es ging uns hier um eine gewisse Anzahl von Geld, sondern es geht hier um (low mumble) Verfahrenswege, um Umgang mit dem Parlament.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness; style: monologue, formal; average recording, no background noise; genuineness 3.7/6; vocal-burst blend 0.7/10; 11.2s, DE.
DE_-OLjkEFs1wk_W000004 · in -20.0 dBFS · gain +0.0 dB · emolia-00247
(steady, no disfluency, clear, formal) Nächste Rednerin ist Kollegin Dorn, Bündnis 90, die Grünen.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: formal, authoritative; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.9/10; 4.1s, DE.
DE_-OLjkEFs1wk_W000007 · in -19.0 dBFS · gain -1.0 dB · emolia-00247
(fairly steady, little disfluency, clear, authoritative) Deshalb werden wir dem Antrag, der übrigens ja nie im Vorfeld diskutiert wurde, dann auch entsprechend ablehnen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, formal; average recording, no background noise; genuineness 2.0/6; vocal-burst blend 0.7/10; 8.0s, DE.
DE_-OLjkEFs1wk_W000010 · in -19.1 dBFS · gain -0.9 dB · emolia-00247
Helplessness rising ↑identity −0.00 emotion 100 %   sad-Helplessness-BASE-k3 · #8

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.07, 0.41, 0.41 — a total move of +0.34.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.34 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 21 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.862 before conversion and 0.858 after — it fell by 0.004. Neighbour-to-neighbour the worst pair went 0.819 → 0.873. (The earlier render, with segment 1 left raw, scores 0.788 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Helplessness moved +0.341 in the original and +0.341 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.82 → 3.05 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0047 → -0.0046 → -0.0044normalised 0.118 → 0.196 → 0.391identity cos to seg 1 0.862 → 0.858 -0.004identity cos neighbours 0.819 → 0.873d_b rescored +0.341 → +0.341d_a rescored +0.341 → +0.341d_a mined 0.000d_b mined 0.273min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-QLrfFrKsMItotal 20.6schain gain +2.0 dBseam step 0.2 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(formal, newsreading) In der Fabrik laufen die Vorbereitungen für die Kampagne, also die Zeit von Ernte und Verarbeitung der Rüben. Sie dauert von Mitte September bis Mitte Dezember.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.1/10; 9.0s, DE.
DE_-QLrfFrKsMI_W000004 · in -20.6 dBFS · gain +0.6 dB · emolia-00147
(formal, newsreading) Beide Maschinen der Becherlinie füllen pro Stunde 6 bis 7000 Becher ab, die in Kartons verpackt und dann palettiert werden.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.1/10; 7.3s, DE.
DE_-QLrfFrKsMI_W000106 · in -18.0 dBFS · gain -2.0 dB · emolia-00147
(formal, newsreading) aber auch die Brotindustrie und pharmazeutische Betriebe verarbeiten das Produkt weiter.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.4/10; 4.7s, DE.
DE_-QLrfFrKsMI_W000109 · in -20.8 dBFS · gain +0.8 dB · emolia-00147
Helplessness rising ↑identity +0.00 emotion 100 %   sad-Helplessness-BASE-k3 · #9

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.41, 0.41, 0.76 — a total move of +0.35.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.35 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 22 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.649 before conversion and 0.649 after — it rose by 0.000. Neighbour-to-neighbour the worst pair went 0.588 → 0.472. (The earlier render, with segment 1 left raw, scores 0.453 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Helplessness moved +0.353 in the original and +0.353 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.60 → 2.92 (+0.32) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0044 → -0.0042 → -0.0039normalised 0.391 → 0.598 → 0.831identity cos to seg 1 0.649 → 0.649 +0.000identity cos neighbours 0.588 → 0.472d_b rescored +0.353 → +0.353d_a rescored +0.353 → +0.353d_a mined 0.001d_b mined 0.439min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-U2hZsyyyRutotal 20.8schain gain +3.8 dBseam step 1.5 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(normal-paced, no disfluency, clear, authoritative) Zunächst mal nehme ich dafür gestrichenes Papier.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 0.0/10; 3.3s, DE.
DE_-U2hZsyyyRu_W000001 · in -17.7 dBFS · gain -2.3 dB · emolia-00147
(infatuation · measured, frequent disfluency, average clarity, casual) (low mumble) euh, puhst du das jetzt mal aus Papier? Man muss ziemlich nah ran an dieses Röhrchen.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as infatuation; style: casual, conversational; good recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.2/10; 4.2s, DE.
DE_-U2hZsyyyRu_W000020 · in -17.1 dBFS · gain -2.9 dB · emolia-00147
(emotional numbness, fatigue exhaustion, disappointment · normal-paced, some disfluency, average clarity, casual) Ich habe da eine Batterie eingelegt, aber das Ganze ist eine Fehlkonstruktion. Hier sind so Stege in der Abdeckung und es lässt sich dann jetzt hier nicht mehr drüber schieben und befestigen. Also als Kinderspielzeug ungeeignet, aber
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as emotional numbness, fatigue exhaustion, disappointment; style: casual, monologue; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 0.0/10; 13.8s, DE.
DE_-U2hZsyyyRu_W000023 · in -17.9 dBFS · gain -2.1 dB · emolia-00147
Helplessness rising ↑identity +0.08 emotion REVERSED   sad-Helplessness-BASE-k3 · #10

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.07, 0.41, 0.41 — a total move of +0.34.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.34 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 29 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.751 before conversion and 0.835 after — it rose by 0.084. Neighbour-to-neighbour the worst pair went 0.842 → 0.866. (The earlier render, with segment 1 left raw, scores 0.761 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Helplessness moved +0.341 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.98 → 3.23 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0047 → -0.0045 → -0.0044normalised 0.118 → 0.289 → 0.391identity cos to seg 1 0.751 → 0.835 +0.084identity cos neighbours 0.842 → 0.866d_b rescored +0.341 → +0.000d_a rescored +0.341 → +0.000d_a mined 0.000d_b mined 0.273min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-WhG-j-UudUtotal 28.3schain gain +3.0 dBseam step 0.8 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(measured, little disfluency, average clarity, authoritative) Hello. Herzlich willkommen aus der Quantum Storm Star Wars Collection. Mein Name ist Deniz und ich möchte euch heute den
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, didactic; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 0.1/10; 6.7s, DE.
DE_-WhG-j-UudU_W000000 · in -19.3 dBFS · gain -0.7 dB · emolia-00238
(disgust · normal-paced, little disfluency, clear, didactic) Und diese Tasten haben ein Soundmodul aktiviert, was auf Basis eines kleinen Plattenspielers basiert.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust; style: didactic, monologue; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.2/10; 7.1s, DE.
DE_-WhG-j-UudU_W000045 · in -19.8 dBFS · gain -0.2 dB · emolia-00238
(relief, concentration · measured, some disfluency, average clarity, monologue) Aufkleber drauf sind. Das finde ich zwar jetzt nicht ganz so schlimm, aber dafür finde ich es umso schöner, dass an dem neuen Truppentransporter keine Aufkleber drauf sind, weil, wie gesagt, die haben natürlich auch immer die Eigenschaft, über die Jahre mal zu verschwinden.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as relief, concentration; style: monologue, casual; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 1.3/10; 14.9s, DE.
DE_-WhG-j-UudU_W000051 · in -19.7 dBFS · gain -0.3 dB · emolia-00238
Helplessness rising ↑identity +0.01 emotion —   sad-Helplessness-BASE-k3 · #11

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.41, 0.41, 0.41 — a total move of +0.00.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.00 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 36 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.822 before conversion and 0.836 after — it rose by 0.013. Neighbour-to-neighbour the worst pair went 0.833 → 0.855. (The earlier render, with segment 1 left raw, scores 0.765 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The re-scored move on Helplessness is +0.000 before and +0.000 after, but the before-value is too close to zero for a retention ratio to mean anything on this chain.

Quality. Mean predicted overall quality across the segments went 2.78 → 3.16 (+0.39) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0045 → -0.0044 → -0.0042normalised 0.289 → 0.391 → 0.598identity cos to seg 1 0.822 → 0.836 +0.013identity cos neighbours 0.833 → 0.855d_b rescored +0.000 → +0.000d_a rescored +0.000 → +0.000d_a mined 0.000d_b mined 0.310min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-YzjWuJuZL8total 35.2schain gain +2.2 dBseam step 0.8 dBcrossfades 100/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, fairly steady
(disappointment, sadness, distress · slightly relaxed, didactic, monologue) (ahem) Und mein Partner, der im ersten Stich noch Kreuz Lusche abgeworfen hat, wechselt jetzt irgendwie ein bisschen inkonsequent die Farbe und wirft jetzt plötzlich eine Pik Lusche ab.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, sadness, distress; style: didactic, monologue; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.0/10; 9.2s, DE.
DE_-YzjWuJuZL8_W000009 · in -17.7 dBFS · gain -2.3 dB · emolia-00238
(slightly relaxed, monologue, casual) Ganz interessant, die Variante, die für die Gegenkartei zu noch einem Auge mehr führt, ist tatsächlich die,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.1/10; 6.4s, DE.
DE_-YzjWuJuZL8_W000027 · in -16.8 dBFS · gain -3.2 dB · emolia-00238
(relief, affection, hope enthusiasm optimism · neutral tension, casual, monologue) Nützen tut's auch nicht. Spaß am Skat hab ich trotzdem nicht verloren. War aber, das muss ich jetzt sagen, tatsächlich richtig eine Therapiesitzung hier heute für mich. Damit bin ich jetzt auch durch. Hoffe, ihr hattet euren Spaß. Hoffe, wir sehen uns bald mit weiteren interessanten Verteilungen. Wünsche euch noch einen hervorragenden Rest Wochenende und eine schöne Woche. Bis bald. Gut Blatt.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, affection, hope enthusiasm optimism; style: casual, monologue; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 3.7/10; 20.0s, DE.
DE_-YzjWuJuZL8_W000039 · in -18.2 dBFS · gain -1.8 dB · emolia-00238
Helplessness rising ↑identity +0.01 emotion —   sad-Helplessness-BASE-k3 · #12

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.41, 0.41, 0.41 — a total move of +0.00.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.00 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 48 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.818 before conversion and 0.824 after — it rose by 0.006. Neighbour-to-neighbour the worst pair went 0.818 → 0.855. (The earlier render, with segment 1 left raw, scores 0.761 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The re-scored move on Helplessness is +0.000 before and +0.000 after, but the before-value is too close to zero for a retention ratio to mean anything on this chain.

Quality. Mean predicted overall quality across the segments went 3.17 → 3.42 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0045 → -0.0043 → -0.0042normalised 0.289 → 0.497 → 0.598identity cos to seg 1 0.818 → 0.824 +0.006identity cos neighbours 0.818 → 0.855d_b rescored +0.000 → +0.000d_a rescored +0.000 → +0.000d_a mined 0.000d_b mined 0.310min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-bnJC1eLO5Itotal 47.6schain gain +2.5 dBseam step 0.7 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, quiet background, slightly relaxed, fairly steady, some disfluency, average clarity
(normal-paced, normally alert, storytelling, monologue) Wir machen eine Geisbergrunde heute miteinander. Man kann um den ganzen Geisberg rum spazieren. Und das Video, das werden wir heute rund um den Geisberg machen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: storytelling, monologue; good recording, quiet background; genuineness 1.9/6; vocal-burst blend 1.0/10; 9.6s, DE.
DE_-bnJC1eLO5I_W000000 · in -18.2 dBFS · gain -1.8 dB · emolia-00113
(bitterness, anger, concentration · normal-paced, normally alert, monologue, storytelling) Also bleiben wir hellwach und versuchen wir so einfach. Und wenn es dir einmal nicht gelingt, macht es überhaupt nichts. Beim nächsten Mal wird es dir besser gelingen. Und beim übernächsten Mal wird es dir aus dem Stand heraus sofort gelingen. Und du wirst merken, auf einmal gehst du mit unvorhergesehenen Situationen viel lockerer, viel leichter um.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as bitterness, anger, concentration; style: monologue, storytelling; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 3.7/10; 19.7s, DE.
DE_-bnJC1eLO5I_W000028 · in -18.4 dBFS · gain -1.6 dB · emolia-00113
(relief, contemplation, pride · measured, subdued, monologue, casual) Und wenn nachher schaust, die Dinge regeln sich oft ganz für selber, oder du warst dem Moment schon was zu tun ist, oder du warst es zu einem späteren Zeitpunkt. Und im Nachhinein sagst du immer, wa, eigentlich ist ganz gut, dass das passiert ist, weil sonst hätte das andere, viel Größere, nicht passieren können.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, contemplation, pride; style: monologue, casual; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 3.5/10; 18.7s, DE.
DE_-bnJC1eLO5I_W000029 · in -18.5 dBFS · gain -1.5 dB · emolia-00113
Helplessness rising ↑identity +0.59 emotion 151 %   sad-Helplessness-BASE-k3 · #13

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.41, 0.41, 0.41 — a total move of +0.00.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.00 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 36 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.221 before conversion and 0.808 after — it rose by 0.587. Neighbour-to-neighbour the worst pair went 0.157 → 0.784. (The earlier render, with segment 1 left raw, scores 0.768 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Helplessness moved +0.353 in the original and +0.532 after conversion — 151 % of the delta retained, i.e. the move came out slightly larger after conversion than before.

Quality. Mean predicted overall quality across the segments went 3.10 → 3.29 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0044 → -0.0043 → -0.0041normalised 0.391 → 0.497 → 0.691identity cos to seg 1 0.221 → 0.808 +0.587identity cos neighbours 0.157 → 0.784d_b rescored +0.353 → +0.532d_a rescored +0.353 → +0.532d_a mined 0.000d_b mined 0.300min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-ffmw_U0FWYtotal 35.3schain gain +1.4 dBseam step 0.3 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, almost no disfluency
(fairly steady, formal, monologue) Die Experten schlagen vor, Sondervermögen in den Haushalt zu integrieren, Steuern und Abgaben auf Arbeit zu senken, digitale Verwaltung zu stärken, zum Beispiel Bürokratie abzubauen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.4/10; 10.1s, DE.
DE_-ffmw_U0FWY_W000003 · in -16.1 dBFS · gain -3.9 dB · emolia-00147
(pride · steady, newsreading, formal) Der Wohlstand der Zukunft wird durch die Dekarbonisierung geschaffen. Klimaschutz schafft Wohlstand, auch industriellen Wohlstand und dem fühlen wir uns verpflichtet.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride; style: newsreading, formal; good recording, quiet background; genuineness 0.9/6; vocal-burst blend 0.5/10; 7.8s, DE.
DE_-ffmw_U0FWY_W000005 · in -17.4 dBFS · gain -2.6 dB · emolia-00147
(disappointment, concentration, anger · fairly steady, formal, monologue) Die Wirtschaft wachse, obwohl die Emissionen sinken, hebt die OECD hervor. Aber noch immer kommen rund drei Viertel des Energieaufkommens aus frustilen Quellen. Nötig sei der Ausbau von E-Mobilität und Schiene, Aufforstung von Wäldern, Renatuierung der Moore und der Schutz vor Auswirkungen des Klimawandels.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, concentration, anger; style: formal, monologue; average recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 17.7s, DE.
DE_-ffmw_U0FWY_W000006 · in -14.4 dBFS · gain -5.6 dB · emolia-00147
Helplessness rising ↑identity +0.05 emotion —   sad-Helplessness-BASE-k3 · #14

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.41, 0.41, 0.41 — a total move of +0.00.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.00 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 41 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.818 before conversion and 0.866 after — it rose by 0.049. Neighbour-to-neighbour the worst pair went 0.871 → 0.897. (The earlier render, with segment 1 left raw, scores 0.739 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The re-scored move on Helplessness is +0.000 before and +0.353 after, but the before-value is too close to zero for a retention ratio to mean anything on this chain.

Quality. Mean predicted overall quality across the segments went 3.02 → 3.29 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0046 → -0.0045 → -0.0043normalised 0.196 → 0.289 → 0.497identity cos to seg 1 0.818 → 0.866 +0.049identity cos neighbours 0.871 → 0.897d_b rescored +0.000 → +0.353d_a rescored +0.000 → +0.353d_a mined 0.000d_b mined 0.301min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-iFXnOKFbJktotal 40.0schain gain +1.3 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, no background noise, normal-paced, normally alert
(contemplation, pride · monologue, casual) Der Hochwassersensor reicht laut Hersteller drei bis fünf Jahre, je nachdem, wie lang man oder wie oft man misst. Also wir messen dort momentan alle zehn Minuten den Abstand vom Sensor zum Boden und liegen dabei, ja, so dreieinhalb Jahre, denke ich mal, da müssen wir die Batterie wechseln, (low mumble) was echt gut ist, weil, (low mumble)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, pride; style: monologue, casual; average recording, no background noise; genuineness 4.7/6; vocal-burst blend 0.0/10; 19.0s, DE.
DE_-iFXnOKFbJk_W000020 · in -24.2 dBFS · gain +4.2 dB · emolia-00095
(monologue, casual) (low mumble) Und falls es mal Hochwasser geben würde, würde man das hier sehen, dann wären die Kästchen rot, (low mumble) auch mit einer Angabe, wie viel Hochwasser wir haben. Also, wir verdienen 20 cm, 30 cm Hochwasser. Und man weiß dann auch, dass es,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; average recording, no background noise; genuineness 4.0/6; vocal-burst blend 0.0/10; 13.5s, DE.
DE_-iFXnOKFbJk_W000076 · in -21.8 dBFS · gain +1.8 dB · emolia-00095
(monologue, casual) Das zu unserem Hochwasserprojekt, das wir mit HTL umgesetzt haben. Wir haben eine kleine Webseite, falls euch das Thema interessiert, lora.olm-digital.com.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; average recording, no background noise; genuineness 2.8/6; vocal-burst blend 0.0/10; 7.9s, DE.
DE_-iFXnOKFbJk_W000078 · in -21.2 dBFS · gain +1.2 dB · emolia-00095
Helplessness rising ↑identity +0.42 emotion —   sad-Helplessness-BASE-k3 · #15

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.41, 0.41, 0.41 — a total move of +0.00.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.00 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 50 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.237 before conversion and 0.656 after — it rose by 0.418. Neighbour-to-neighbour the worst pair went 0.280 → 0.484. (The earlier render, with segment 1 left raw, scores 0.545 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

The re-scored move on Helplessness is +0.000 before and -0.341 after, but the before-value is too close to zero for a retention ratio to mean anything on this chain.

Quality. Mean predicted overall quality across the segments went 2.77 → 3.16 (+0.39) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0046 → -0.0045 → -0.0043normalised 0.196 → 0.289 → 0.497identity cos to seg 1 0.237 → 0.656 +0.418identity cos neighbours 0.280 → 0.484d_b rescored +0.000 → -0.341d_a rescored +0.000 → -0.341d_a mined 0.000d_b mined 0.301min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-xg8lm69_K8total 48.9schain gain +3.0 dBseam step 3.2 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-bright, fairly smooth, balanced body, average recording
(concentration, pain · measured, normally alert, slightly relaxed, newsreading) Gesetz zur Stärkung der Wahlbeteiligung bei Gremienwahlen an hessischen Hochschulen, Drucksache 20.3998. Zur Einbringung hat sich Herr Dr. Büger von der FDP gemeldet. Die vereinbarte Redezeit beträgt fünf Minuten.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, fairly guarded; reads as concentration, pain; style: newsreading, formal; average recording, quiet background; genuineness 0.5/6; vocal-burst blend 0.5/10; 16.3s, DE.
DE_-xg8lm69_K8_W000002 · in -25.5 dBFS · gain +5.5 dB · emolia-00224
(normal-paced, normally alert, slightly relaxed, formal) Für die Fraktion der LINKEN hat sich ihre Vorsitzende, Frau Wissler, zu Wort gemeldet.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.7/10; 4.8s, DE.
DE_-xg8lm69_K8_W000048 · in -22.1 dBFS · gain +2.1 dB · emolia-00224
(sourness, doubt, interest · brisk, energised, neutral tension, ranting) In dem Fall zur verfassten Studierendenschaft. Dabei ist ja eine Diskussion, die wir generell (ahem) auch führen, wie wir jetzt unter Pandemiebedingungen, wo wir gerade Schwierigkeiten haben, (ahem) Listen aufzustellen, Wahlversammlungen zu machen. (ahem) Da haben wir an der Stelle die CDU und DIE LINKE eine Gemeinsamkeit insofern, dass wir schon das ganze Jahr versuchen, Vorsitzende zu wählen. Sie schaffen es nicht, wir schaffen es nicht, weil man unter Pandemie-Gesichtspunkten gerade schwierig Parteitage machen kann und elektronische Wahlen nicht möglich sind. (low mumble) (ahem)
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, fairly guarded; reads as sourness, doubt, interest; style: ranting, dramatic; average recording, some background noise; genuineness 2.9/6; vocal-burst blend 4.5/10; 28.3s, DE.
DE_-xg8lm69_K8_W000049 · in -22.4 dBFS · gain +2.4 dB · emolia-00224
Helplessness rising ↑identity −0.06 emotion REVERSED   sad-Helplessness-BASE-k3 · #16

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.41, 0.41, 0.76 — a total move of +0.35.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.35 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 15 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.755 before conversion and 0.699 after — it fell by 0.056. Neighbour-to-neighbour the worst pair went 0.755 → 0.691. (The earlier render, with segment 1 left raw, scores 0.703 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Helplessness moved +0.353 in the original and -0.119 after conversion — it changed direction. On this chain the corrected audio is not an improvement.

Quality. Mean predicted overall quality across the segments went 2.68 → 2.97 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0043 → -0.0041 → -0.0039normalised 0.497 → 0.691 → 0.831identity cos to seg 1 0.755 → 0.699 -0.056identity cos neighbours 0.755 → 0.691d_b rescored +0.353 → -0.119d_a rescored +0.353 → -0.119d_a mined 0.000d_b mined 0.334min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_-xpBDPzvHIItotal 14.5schain gain +1.7 dBseam step 0.5 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, balanced body, good recording
(doubt · measured, normally alert, slightly relaxed, formal) Du kennst bestimmt so die Situation. Du hast schon lange irgendwas nicht mehr gegessen.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as doubt; style: formal, didactic; good recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.0/10; 5.3s, DE.
DE_-xpBDPzvHII_W000000 · in -18.9 dBFS · gain -1.1 dB · emolia-00182
(longing · normal-paced, normally alert, slightly relaxed, conversational) Let's go! Jetzt, die letzten Wochen, jetzt in diesem Jahr hier noch.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as longing; style: conversational, casual; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 1.2/10; 4.2s, DE.
DE_-xpBDPzvHII_W000028 · in -15.7 dBFS · gain -4.3 dB · emolia-00182
(thankfulness gratitude, contempt, affection · measured, very low-energy, relaxed, ASMR) Das kannst nur du dich, für dein Leben, für deine Freiheit, für deine Leichtigkeit.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; clear, no disfluency, fairly narrow pitch, audible breath; affect is mildly positive, neutral stance, neutral openness; reads as thankfulness gratitude, contempt, affection; style: ASMR, whispered; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 2.7/10; 5.5s, DE.
DE_-xpBDPzvHII_W000029 · in -19.9 dBFS · gain -0.1 dB · emolia-00182
Helplessness rising ↑identity +0.51 emotion —   sad-Helplessness-BASE-k3 · #17

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.41, 0.41, 0.41 — a total move of +0.00.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.00 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 22 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.034 before conversion and 0.543 after — it rose by 0.509. Neighbour-to-neighbour the worst pair went 0.034 → 0.522. (The earlier render, with segment 1 left raw, scores 0.467 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

The re-scored move on Helplessness is +0.000 before and +0.353 after, but the before-value is too close to zero for a retention ratio to mean anything on this chain.

Quality. Mean predicted overall quality across the segments went 2.57 → 2.99 (+0.42) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0044 → -0.0042 → -0.0041normalised 0.391 → 0.598 → 0.691identity cos to seg 1 0.034 → 0.543 +0.509identity cos neighbours 0.034 → 0.522d_b rescored +0.000 → +0.353d_a rescored +0.000 → +0.353d_a mined 0.000d_b mined 0.300min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_02FqQjJyMtytotal 21.4schain gain +3.5 dBseam step 1.2 dBcrossfades 100/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, quiet background, normally alert, slightly relaxed, fairly steady
(measured, some disfluency, somewhat unclear, casual) (low mumble) euh, wie machen wir uns hier, wir sprechen jetzt Sebastian, Sebastian Kruppter. Das ist unsere nächste Quest. Das machen wir einfach jetzt in der Aufnahme weiter.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 1.6/10; 8.6s, DE.
DE_02FqQjJyMty_W000003 · in -22.4 dBFS · gain +2.4 dB · emolia-00215
(jealousy and envy, sourness, malevolence malice · normal-paced, almost no disfluency, clear, storytelling) Sie hat schnell begriffen, als du ihr gezeigt hast, wie man aus dem Zelt entkommt. Sobald sie das Ei sieht, wird sie verstehen warum wir hier sind. Dann können wir die Wilderer ein für alle mal erledigen.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as jealousy and envy, sourness, malevolence malice; style: storytelling, narration; good recording, quiet background; genuineness 0.6/6; vocal-burst blend 0.2/10; 9.9s, DE.
DE_02FqQjJyMty_W000029 · in -17.7 dBFS · gain -2.3 dB · emolia-00215
(emotional numbness, fatigue exhaustion · measured, some disfluency, slurred, casual) Bist du was? Ich werd's mal in welchen Tragung mal schauen.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as emotional numbness, fatigue exhaustion; style: casual, conversational; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 0.0/10; 3.4s, DE.
DE_02FqQjJyMty_W000043 · in -20.9 dBFS · gain +0.9 dB · emolia-00215
Helplessness rising ↑identity +0.13 emotion REVERSED   sad-Helplessness-BASE-k3 · #18

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.41, 0.41, 0.76 — a total move of +0.35.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.35 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 18 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.579 before conversion and 0.713 after — it rose by 0.134. Neighbour-to-neighbour the worst pair went 0.579 → 0.748. (The earlier render, with segment 1 left raw, scores 0.645 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Helplessness moved +0.353 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.85 → 3.03 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0042 → -0.0041 → -0.0038normalised 0.598 → 0.691 → 0.862identity cos to seg 1 0.579 → 0.713 +0.134identity cos neighbours 0.579 → 0.748d_b rescored +0.353 → +0.000d_a rescored +0.353 → +0.000d_a mined 0.000d_b mined 0.264min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_03AB_mR589Qtotal 17.7schain gain +2.2 dBseam step 2.1 dBcrossfades 150/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, balanced body, average recording, quiet background, measured, slightly relaxed, fairly steady, frequent disfluency
(very low-energy, somewhat unclear, fairly narrow pitch, monologue) Ja. Und wie wir schon gelernt haben, es gibt so, (ahem) äh, Größen, die halt so.
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: monologue, ASMR; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 1.1/10; 5.7s, DE.
DE_03AB_mR589Q_W000001 · in -16.1 dBFS · gain -3.9 dB · emolia-00003
(normally alert, slurred, fairly narrow pitch, whispered) Vektor mehr, sondern eine Zahl, ein sogenanntes Skalar. (low mumble) Und das war halt,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: whispered, monologue; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.0/10; 5.7s, DE.
DE_03AB_mR589Q_W000005 · in -17.8 dBFS · gain -2.2 dB · emolia-00003
(doubt, contemplation · subdued, somewhat unclear, moderate pitch range, monologue) Es kann sein, dass, (low mumble) wenn jemand hier das versucht, also diese Errechnung zu machen, dass hier eine andere Zahl vorkommt.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as doubt, contemplation; style: monologue, casual; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 0.0/10; 6.8s, DE.
DE_03AB_mR589Q_W000010 · in -18.8 dBFS · gain -1.2 dB · emolia-00003
Helplessness rising ↑identity +0.02 emotion 100 %   sad-Helplessness-BASE-k3 · #19

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.41, 0.41, 0.76 — a total move of +0.35.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.35 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 26 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.747 before conversion and 0.762 after — it rose by 0.015. Neighbour-to-neighbour the worst pair went 0.747 → 0.817. (The earlier render, with segment 1 left raw, scores 0.788 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Helplessness moved +0.353 in the original and +0.353 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 3.13 → 3.26 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0043 → -0.0041 → -0.0038normalised 0.497 → 0.691 → 0.862identity cos to seg 1 0.747 → 0.762 +0.015identity cos neighbours 0.747 → 0.817d_b rescored +0.353 → +0.353d_a rescored +0.353 → +0.353d_a mined 0.001d_b mined 0.365min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0IWgiqWfPNAtotal 24.9schain gain +2.0 dBseam step 0.9 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, some disfluency
(slightly relaxed, moderately variable, wide pitch range, conversational) Hey Leute, willkommen zurück zu einem neuen Video hier auf Popel mit Zucker. Wie immer fangen wir an mit einem extra Schluck.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: conversational, playful; good recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.4/10; 5.6s, DE.
DE_0IWgiqWfPNA_W000000 · in -17.9 dBFS · gain -2.1 dB · emolia-00195
(confusion · slightly relaxed, fairly steady, moderate pitch range, monologue) Die Übergabe der, der, der Geiseln, das war einfach zu einfach. Die hätten nicht einfach nen Andes dahingestellt und gesagt, jo, das ist nen Andes. Das.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion; style: monologue, casual; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 0.0/10; 10.9s, DE.
DE_0IWgiqWfPNA_W000052 · in -18.4 dBFS · gain -1.6 dB · emolia-00195
(contemplation, relief, fatigue exhaustion · relaxed, fairly steady, moderate pitch range, casual) Meiner Meinung nach war das ein riesiger Fehleinsatz deswegen und, ja. Verständlich, dass dann am Ende das Resultat so rauskommt. Ja.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as contemplation, relief, fatigue exhaustion; style: casual, monologue; average recording, quiet background; mildly explicit content; genuineness 4.4/6; vocal-burst blend 1.4/10; 8.8s, DE.
DE_0IWgiqWfPNA_W000060 · in -21.6 dBFS · gain +1.6 dB · emolia-00195
Helplessness rising ↑identity +0.02 emotion 100 %   sad-Helplessness-BASE-k3 · #20

This is not a strict-rule trajectory. It comes from rescue rule BASE, which exists because the strict rule returns nothing at all for Helplessness.

The raw scorer output across the chain is -0.00, -0.00, -0.00. In the first clip the scorer found no Helplessness whatsoever (-0.00); by the last it is at -0.00.

On the corpus-wide percentile scale those become 0.41, 0.41, 0.76 — a total move of +0.35.

That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Helplessness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.35 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.

Re-ranked among the Helplessness-bearing clips only — which is exactly what this rule does — the same clips read 0.00, 0.00, 0.00, a move of +0.00, which is usable again.

What this rule changes: The strict rule, unchanged, as a control. Rank-normalise the emotion over the whole corpus, then require the chain to rise by at least 0.25 end-to-end with every consecutive step at most 0.25.

What it costs: Nothing -- this is the strict rule everything else is measured against.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 33 s · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.837 before conversion and 0.854 after — it rose by 0.017. Neighbour-to-neighbour the worst pair went 0.837 → 0.820. (The earlier render, with segment 1 left raw, scores 0.675 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Helplessness moved +0.353 in the original and +0.353 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.84 → 3.20 (+0.36) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3raw scores -0.0043 → -0.0042 → -0.004normalised 0.497 → 0.598 → 0.770identity cos to seg 1 0.837 → 0.854 +0.017identity cos neighbours 0.837 → 0.820d_b rescored +0.353 → +0.353d_a rescored +0.353 → +0.353d_a mined 0.000d_b mined 0.274min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_0KZXkf63R18total 32.8schain gain +2.6 dBseam step 0.2 dBcrossfades 100/100 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, quiet background, slightly relaxed, moderate pitch range
(affection · normal-paced, normally alert, moderately variable, casual) Ich bin Capriela, psychologische Berater und ja, spezialisiert auf KPTBS. (ahem) Ich leide auch darunter und jetzt freue mich einfach in diesem Video auf Dich.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as affection; style: casual, conversational; good recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.0/10; 12.4s, DE.
DE_0KZXkf63R18_W000002 · in -21.1 dBFS · gain +1.1 dB · emolia-00254
(measured, very low-energy, fairly steady, monologue) (wistful sigh) Dann sollte das mit dem Therapeut trocken werden, weil eine Stabilisierung, also ein, ein nochmals erhöhte Stabilisierung ist dann notwendig.
full caption & clip details
A young adult somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 0.0/10; 10.3s, DE.
DE_0KZXkf63R18_W000056 · in -17.3 dBFS · gain -2.7 dB · emolia-00254
(relief, hope enthusiasm optimism, elation · measured, normally alert, fairly steady, casual) Ich hoffe, dass ich auch diesmal, (low mumble) ähm, einige Anregungen, Inspiration und ja, Empfehlungen aussprechen konnte und dass es dir hilfreich war.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, audible breath; affect is mildly positive, slightly submissive, neutral openness; reads as relief, hope enthusiasm optimism, elation; style: casual, monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 0.0/10; 10.3s, DE.
DE_0KZXkf63R18_W000057 · in -20.5 dBFS · gain +0.5 dB · emolia-00254