B1: 2 chains from each of the 10 scarcest ordered emotion pairs (supply 8-21 chains each).
This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_rare-B1-pairs.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words
are a real non-speech sound, printed where it happens.
The full generated caption for any clip is under “full caption & clip
details”. Its perceived-gender and background-noise clauses were re-rendered from
the numeric buckets, because the versions stored in the corpus index had those two
ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
The tags and captions describe the ORIGINAL clips — they are the
corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the
converted audio are the ones printed in each card's own paragraph and chips.
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Sexual Lust barely moves at all, sitting near 1.00 throughout.
It takes 2 clips to get there. Clip to clip the moves are +0.20 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.72 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.72 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 14 s · en · emolia
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.544 before conversion and 0.739 after — it rose by 0.195. Neighbour-to-neighbour the worst pair went 0.544 → 0.739. (The earlier render, with segment 1 left raw, scores 0.653 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.205 in the original and +0.199 after conversion — 97 % of the delta retained, which is essentially all of it. On the other named axis, Sexual Lust, -0.043 became -0.064.
Quality. Mean predicted overall quality across the segments went 2.81 → 2.92 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.544 → 0.739+0.195identity cos neighbours 0.544 → 0.739d_b rescored +0.205 → +0.199d_a rescored -0.043 → -0.064d_a mined -0.043d_b mined 0.205min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00057_S02002total 14.2schain gain +4.4 dBseam step 0.2 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a child feminine voice · no background noise, brisk, almost no disfluency, wide pitch range
(sexual lust, infatuation, disgust · energised, neutral tension, moderately variable, storytelling)So how do you want me to react to it? When you flirt with a girl in front of my face? Should I say congrats to the two of you?
full caption & clip details
A child feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as sexual lust, infatuation, disgust; style: storytelling, conversational; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.3/10; 8.8s, EN.
EN_B00057_S02002_W000026 · in -15.3 dBFS · gain -4.7 dB · emolia-01338
(disappointment, infatuation, distress·highly aroused, slightly tense, variable, storytelling)I can't stand you anymore. You make me look like a dump. I think we should break up.
full caption & clip details
A young adult feminine voice; delivery is highly aroused, brisk, slightly tense, variable; timbre is slightly cool, neutral-bright, very rough, thin; average clarity, almost no disfluency, wide pitch range, audible breath; affect is negative, slightly dominant, fairly guarded; reads as disappointment, infatuation, distress; style: storytelling, casual; average recording, no background noise; mildly explicit content; genuineness 2.0/6; vocal-burst blend 0.4/10; 5.5s, EN.
EN_B00057_S02002_W000027 · in -16.3 dBFS · gain -3.7 dB · emolia-01338
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Sexual Lust drifts down from 0.91 to 0.81 (-0.10), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.06, then +0.11, then +0.04 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.76 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.76 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 48 s · ja · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.796 before conversion and 0.861 after — it rose by 0.065. Neighbour-to-neighbour the worst pair went 0.758 → 0.790. (The earlier render, with segment 1 left raw, scores 0.796 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.200 in the original and +0.135 after conversion — 68 % of the delta retained. On the other named axis, Sexual Lust, -0.096 became -0.024.
Quality. Mean predicted overall quality across the segments went 2.85 → 3.15 (+0.31) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.796 → 0.861+0.065identity cos neighbours 0.758 → 0.790d_b rescored +0.200 → +0.135d_a rescored -0.096 → -0.024d_a mined -0.097d_b mined 0.200min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang jaspeaker JA_sOi86mTaj0ytotal 47.2schain gain +0.6 dBseam step 1.1 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-bright, quiet background, fast, neutral tension, moderately variable, some disfluency, average clarity, light breath
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Helplessness barely moves at all, sitting near 1.00 throughout.
It takes 5 clips to get there. Clip to clip the moves are +0.00, then +0.00, then +0.00, then +0.20 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 75 s · spanish · mls
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.965 before conversion and 0.941 after — it fell by 0.024. Neighbour-to-neighbour the worst pair went 0.965 → 0.949. (The earlier render, with segment 1 left raw, scores 0.803 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.201 in the original and +0.502 after conversion — 250 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Helplessness, -0.041 became -0.107.
Quality. Mean predicted overall quality across the segments went 3.19 → 3.35 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.965 → 0.941-0.024identity cos neighbours 0.965 → 0.949d_b rescored +0.201 → +0.502d_a rescored -0.041 → -0.107d_a mined -0.040d_b mined 0.202min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang spanishspeaker 3946total 74.0schain gain -1.5 dBseam step 0.7 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert, slightly relaxed, fairly steady
(helplessness, fatigue exhaustion, distress · normal-paced, some disfluency, somewhat unclear, monologue)el tiempo es breve las ansias crecen las esperanzas menguan y con todo esto llevo la vida sobre el deseo que tengo de vivir y quisiera yo no ponerle coto hasta besar los piés vuesa excelencia
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness, fatigue exhaustion, distress; style: monologue, authoritative; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 5.6/10; 12.9s, SPANISH.
3946_10158_000730 · in -20.0 dBFS · gain -0.0 dB · mls-00023
(helplessness, fatigue exhaustion, distress · normal-paced, some disfluency, somewhat unclear, monologue)el tiempo es breve las ansias crecen las esperanzas menguan y con todo esto llevo la vida sobre el deseo que tengo de vivir y quisiera yo no ponerle coto hasta besar los piés vuesa excelencia
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness, fatigue exhaustion, distress; style: monologue, authoritative; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 5.6/10; 12.9s, SPANISH.
3946_10158_000730 · in -20.0 dBFS · gain -0.0 dB · mls-00023
(disgust, helplessness, longing· normal-paced, almost no disfluency, clear, monologue)que podria ser fuese tanto el contento de ver á vuesa excelencia bueno en españa que me volviese á dar la vida pero si está decretado que la haya de perder cúmplase la voluntad de los cielos y por lo ménos sepa vuesa excelencia este mi deseo
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, helplessness, longing; style: monologue, narration; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 3.0/10; 16.1s, SPANISH.
3946_10158_001210 · in -19.3 dBFS · gain -0.7 dB · mls-00024
(disgust, helplessness, longing · normal-paced, almost no disfluency, clear, monologue)que podria ser fuese tanto el contento de ver á vuesa excelencia bueno en españa que me volviese á dar la vida pero si está decretado que la haya de perder cúmplase la voluntad de los cielos y por lo ménos sepa vuesa excelencia este mi deseo
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, helplessness, longing; style: monologue, narration; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 3.0/10; 16.1s, SPANISH.
3946_10158_001210 · in -19.3 dBFS · gain -0.7 dB · mls-00024
(disappointment, shame, bitterness·measured, little disfluency, somewhat unclear, monologue)y sepa que tuvo en mí un tan aficionado criado de servirle que quiso pasar áun más allá de la muerte mostrando su intencion con todo esto como en profecía me alegro de la llegada de vuesa excelencia regocijome de verle señalar con el dedo y
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, shame, bitterness; style: monologue, narration; average recording, quiet background; genuineness 0.7/6; vocal-burst blend 3.2/10; 16.8s, SPANISH.
3946_10158_000604 · in -19.2 dBFS · gain -0.8 dB · mls-00023
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Helplessness barely moves at all, sitting near 1.00 throughout.
It takes 5 clips to get there. Clip to clip the moves are +0.00, then +0.00, then +0.20, then +0.00 — a plateau around step 2, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 79 s · spanish · mls
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.965 before conversion and 0.950 after — it fell by 0.014. Neighbour-to-neighbour the worst pair went 0.965 → 0.950. (The earlier render, with segment 1 left raw, scores 0.791 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.201 in the original and +0.054 after conversion — 27 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Helplessness, -0.041 became -0.108.
Quality. Mean predicted overall quality across the segments went 3.20 → 3.32 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.965 → 0.950-0.014identity cos neighbours 0.965 → 0.950d_b rescored +0.201 → +0.054d_a rescored -0.041 → -0.108d_a mined -0.041d_b mined 0.201min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang spanishspeaker 3946total 77.9schain gain -1.8 dBseam step 0.7 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert, slightly relaxed, fairly steady
(helplessness, fatigue exhaustion, distress · normal-paced, some disfluency, somewhat unclear, monologue)el tiempo es breve las ansias crecen las esperanzas menguan y con todo esto llevo la vida sobre el deseo que tengo de vivir y quisiera yo no ponerle coto hasta besar los piés vuesa excelencia
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness, fatigue exhaustion, distress; style: monologue, authoritative; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 5.6/10; 12.9s, SPANISH.
3946_10158_000730 · in -20.0 dBFS · gain -0.0 dB · mls-00023
(disgust, helplessness, longing· normal-paced, almost no disfluency, clear, monologue)que podria ser fuese tanto el contento de ver á vuesa excelencia bueno en españa que me volviese á dar la vida pero si está decretado que la haya de perder cúmplase la voluntad de los cielos y por lo ménos sepa vuesa excelencia este mi deseo
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, helplessness, longing; style: monologue, narration; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 3.0/10; 16.1s, SPANISH.
3946_10158_001210 · in -19.3 dBFS · gain -0.7 dB · mls-00024
(disgust, helplessness, longing · normal-paced, almost no disfluency, clear, monologue)que podria ser fuese tanto el contento de ver á vuesa excelencia bueno en españa que me volviese á dar la vida pero si está decretado que la haya de perder cúmplase la voluntad de los cielos y por lo ménos sepa vuesa excelencia este mi deseo
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, helplessness, longing; style: monologue, narration; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 3.0/10; 16.1s, SPANISH.
3946_10158_001210 · in -19.3 dBFS · gain -0.7 dB · mls-00024
(disappointment, shame, bitterness·measured, little disfluency, somewhat unclear, monologue)y sepa que tuvo en mí un tan aficionado criado de servirle que quiso pasar áun más allá de la muerte mostrando su intencion con todo esto como en profecía me alegro de la llegada de vuesa excelencia regocijome de verle señalar con el dedo y
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, shame, bitterness; style: monologue, narration; average recording, quiet background; genuineness 0.7/6; vocal-burst blend 3.2/10; 16.8s, SPANISH.
3946_10158_000604 · in -19.2 dBFS · gain -0.8 dB · mls-00023
(disappointment, shame, bitterness · measured, little disfluency, somewhat unclear, monologue)y sepa que tuvo en mí un tan aficionado criado de servirle que quiso pasar áun más allá de la muerte mostrando su intencion con todo esto como en profecía me alegro de la llegada de vuesa excelencia regocijome de verle señalar con el dedo y
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, shame, bitterness; style: monologue, narration; average recording, quiet background; genuineness 0.7/6; vocal-burst blend 3.2/10; 16.8s, SPANISH.
3946_10158_000604 · in -19.2 dBFS · gain -0.8 dB · mls-00023
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Pleasure Ecstasy drifts down from 0.99 to 0.63 (-0.35), which the rule did not require.
It takes 2 clips to get there. Clip to clip the moves are +0.20 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 28 s · it · eurospeech
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.953 before conversion and 0.950 after — it fell by 0.004. Neighbour-to-neighbour the worst pair went 0.953 → 0.950. (The earlier render, with segment 1 left raw, scores 0.811 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.203 in the original and +0.601 after conversion — 296 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Pleasure Ecstasy, -0.354 became -0.225.
Quality. Mean predicted overall quality across the segments went 2.88 → 3.09 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.953 → 0.950-0.004identity cos neighbours 0.953 → 0.950d_b rescored +0.203 → +0.601d_a rescored -0.354 → -0.225d_a mined -0.354d_b mined 0.203min_cos_consec (site) —min_cos_anchor (site) —dataset eurospeechlang itspeaker italy_16_277total 27.3schain gain -0.0 dBseam step 3.2 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an elderly feminine voice · slightly cool, slightly bright, fairly smooth, thin, energised, neutral tension, moderately variable, some disfluency
(pleasure ecstasy, awe, contentment · brisk, very clear, dramatic, ranting)che negli ultimi dieci anni hanno perso moltissimo potere di acquisto, sia nella direzione di una spinta alla ripresa dei consumi come stimolo indispensabile per
full caption & clip details
An elderly feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; very clear, some disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, guarded; reads as pleasure ecstasy, awe, contentment; style: dramatic, ranting; poor recording, noisy background; genuineness 3.5/6; vocal-burst blend 7.0/10; 12.4s, IT.
italy_16_277_1945888_1958272 · in -16.7 dBFS · gain -3.3 dB · eurospeech-01678
(disappointment, shame, distress·fast, average clarity, dramatic, ranting)avviare il percorso di uscita dalla crisi nel nostro Paese. Questa finanziaria non si sta occupando del lavoro e dei redditi da lavoro e da pensione, e a noi questo sembra gravissimo in un momento in cui
full caption & clip details
A young adult feminine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, guarded; reads as disappointment, shame, distress; style: dramatic, ranting; below-average recording, quiet background; genuineness 3.8/6; vocal-burst blend 8.3/10; 15.2s, IT.
italy_16_277_1958272_1973440 · in -16.2 dBFS · gain -3.8 dB · eurospeech-01678
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Pleasure Ecstasy drifts down from 0.99 to 0.30 (-0.69), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.04, then +0.17 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.62 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.68 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.62, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 62 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.631 before conversion and 0.865 after — it rose by 0.234. Neighbour-to-neighbour the worst pair went 0.670 → 0.865. (The earlier render, with segment 1 left raw, scores 0.532 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.207 in the original and +0.185 after conversion — 90 % of the delta retained, which is most of it. On the other named axis, Pleasure Ecstasy, -0.690 became -0.694.
Quality. Mean predicted overall quality across the segments went 2.98 → 3.17 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.631 → 0.865+0.234identity cos neighbours 0.670 → 0.865d_b rescored +0.207 → +0.185d_a rescored -0.690 → -0.694d_a mined -0.691d_b mined 0.204min_cos_consec (site) 0.6838min_cos_anchor (site) 0.6151dataset podcastlang enspeaker 44235total 61.8schain gain +5.0 dBseam step 0.3 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, neutral tension, average clarity
(pleasure ecstasy, teasing, amusement · normal-paced, energised, moderately variable, casual)Yes. And the and the flip out about the camera, like his camera looked worse than how we had it.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as pleasure ecstasy, teasing, amusement; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.9/6; vocal-burst blend 10.0/10; 16.1s, EN.
44235_00407704 · in -25.6 dBFS · gain +5.6 dB · podcast-05745
(amusement, teasing, confusion· normal-paced, normally alert, moderately variable, casual)okay because he wanted his camera to be a certain way, and that and so people think we did that to him, but we didn't. Like we the here's the thing if someone I don't know what his last (low mumble) um his last outburst was, but it was like I think it was like fuck America or something. I don't care. Remember, you remember exploited had to fuck the USA? I
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as amusement, teasing, confusion; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 6.0/6; vocal-burst blend 9.8/10; 20.2s, EN.
44235_00409784 · in -23.5 dBFS · gain +3.5 dB · podcast-05738
(disappointment, shame, impatience and irritability·slow, energised, variable, casual)It's uh (ahem) so it but that doesn't make me a bad person. I just choose to focus on different things. Someone's saying there's there's literally five million people ahead of Sebastian Bach doing worse things to America that you can focus on.
full caption & clip details
A young adult masculine voice; delivery is energised, slow, neutral tension, variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as disappointment, shame, impatience and irritability; style: casual, dramatic; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 5.9/10; 25.9s, EN.
44235_00411984 · in -21.8 dBFS · gain +1.8 dB · podcast-06281
Disappointment ↑ (unconstrained axis: Intoxication Altered States of Consciousness)identity +0.04emotion 94 % rare-B1-pairs · #7
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Intoxication Altered States of Consciousness drifts down from 0.96 to 0.57 (-0.40), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.08, then +0.12 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 51 s · da · eurospeech
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.914 before conversion and 0.951 after — it rose by 0.037. Neighbour-to-neighbour the worst pair went 0.914 → 0.951. (The earlier render, with segment 1 left raw, scores 0.859 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.203 in the original and +0.192 after conversion — 94 % of the delta retained, which is essentially all of it. On the other named axis, Intoxication Altered States of Consciousness, -0.397 became -0.833.
Quality. Mean predicted overall quality across the segments went 3.20 → 3.45 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.914 → 0.951+0.037identity cos neighbours 0.914 → 0.951d_b rescored +0.203 → +0.192d_a rescored -0.397 → -0.833d_a mined -0.397d_b mined 0.203min_cos_consec (site) —min_cos_anchor (site) —dataset eurospeechlang daspeaker denmark_20191M063_2020-02-total 50.2schain gain +2.1 dBseam step 0.7 dBcrossfades 150/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, neutral tension, light breath
(intoxication altered states of consciousness, confusion, impatience and irritability · normal-paced, normally alert, fairly steady, monologue)(low mumble) den nyere politiske danmarkshistorie har gjort mest for at sikre et Danmark i bedre balance – udflytning af statslige arbejdspladser og en lang række øvrige initiativer. (low mumble) Dem kunne vi bruge rigtig meget tid på. Så (low mumble) at vi i Venstre nærer et dybtfølt ønske om at sikre et Danmark i bedre balance, er jo helt (low mumble) logisk.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness, confusion, impatience and irritability; style: monologue, conversational; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 6.0/10; 19.8s, DA.
denmark_20191M063_2020-02-19_1300_1426623_1446447 · in -22.5 dBFS · gain +2.5 dB · eurospeech-00383
(confusion, jealousy and envy, embarrassment· normal-paced, normally alert, moderately variable, conversational)(low mumble) så gør det måske ikke så meget. For hvis man kigger på indholdet i det, kan man jo tage Finansieringsudvalgets rapport fra 2018. Der var der en overkompensation på udlændingeområdet. Og hvad gør regeringen? Følger den så anbefalingen? Nej, det gør den ikke.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, jealousy and envy, embarrassment; style: conversational, didactic; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 4.4/10; 14.6s, DA.
denmark_20191M063_2020-02-19_1300_1463264_1477840 · in -22.2 dBFS · gain +2.2 dB · eurospeech-00383
(disappointment, sourness, impatience and irritability·brisk, energised, moderately variable, monologue)Så den overkompensation, der er beskrevet tidligere, fastholder regeringen (low mumble) i et betydeligt omfang. Det er da fair nok, for det er politik. Men så læg da tallene fuldstændig åbent frem, så vi kan se, præcis hvilken betydning det har for de enkelte kommuner, i stedet for at der er
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as disappointment, sourness, impatience and irritability; style: monologue, dramatic; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 4.0/10; 16.2s, DA.
denmark_20191M063_2020-02-19_1300_1477840_1494000 · in -22.7 dBFS · gain +2.7 dB · eurospeech-00383
Disappointment ↑ (unconstrained axis: Intoxication Altered States of Consciousness)identity +0.16emotion REVERSED rare-B1-pairs · #8
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.21.
Nothing was asked of the other axis, and in fact Intoxication Altered States of Consciousness drifts down from 0.99 to 0.33 (-0.66), which the rule did not require.
It takes 2 clips to get there. Clip to clip the moves are +0.21 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.67 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.67 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 25 s · de · emolia
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.603 before conversion and 0.758 after — it rose by 0.155. Neighbour-to-neighbour the worst pair went 0.603 → 0.758. (The earlier render, with segment 1 left raw, scores 0.681 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
The emotional move did not survive. Re-scored end to end, Disappointment moved +0.207 in the original and -0.408 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Intoxication Altered States of Consciousness, -0.660 became -0.841.
Quality. Mean predicted overall quality across the segments went 2.90 → 3.12 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.603 → 0.758+0.155identity cos neighbours 0.603 → 0.758d_b rescored +0.207 → -0.408d_a rescored -0.660 → -0.841d_a mined -0.660d_b mined 0.207min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_U3ZCmeVgf_Utotal 24.5schain gain +3.4 dBseam step 1.2 dBcrossfades 100 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · fairly smooth, average recording, quiet background, moderately variable, wide pitch range, normal breath
(intoxication altered states of consciousness, elation, contentment · normal-paced, normally alert, relaxed, casual)Bam, Antigra, warum kann ich Antigra? Das ist keine gute Attacke, oder? Also, ich hab's immer als schlechte Attacke empfunden. Weiß nicht, wie ihr das seht, Leute. Hm, hm, hm, wie lang geht da Feuerberg? Ich glaub, 20 Dinger oder so. Naja, ich werd euch (low mumble) irgendwas reinschneiden, falls irgendwas interessantes passiert.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as intoxication altered states of consciousness, elation, contentment; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.7/6; vocal-burst blend 0.6/10; 19.1s, DE.
DE_U3ZCmeVgf_U_W000044 · in -19.5 dBFS · gain -0.5 dB · emolia-00214
(disappointment, impatience and irritability, anger·fast, highly aroused, slightly tense, ranting)Doch diesmal wird es nicht mehr so laufen. Wir haben uns komplett entwickelt und haben Zapdos dabei.
full caption & clip details
A young adult masculine voice; delivery is highly aroused, fast, slightly tense, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; very clear, no disfluency, wide pitch range, normal breath; affect is elated, very dominant, guarded; reads as disappointment, impatience and irritability, anger; style: ranting, dramatic; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 2.2/10; 5.5s, DE.
DE_U3ZCmeVgf_U_W000045 · in -12.4 dBFS · gain -7.6 dB · emolia-00214
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.21.
Nothing was asked of the other axis, and in fact Amusement drifts down from 1.00 to 0.92 (-0.08), which the rule did not require.
It takes 2 clips to get there. Clip to clip the moves are +0.21 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.38 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.38 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 28 s · en · emolia
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.413 before conversion and 0.496 after — it rose by 0.083. Neighbour-to-neighbour the worst pair went 0.413 → 0.496. (The earlier render, with segment 1 left raw, scores 0.452 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.206 in the original and +0.581 after conversion — 283 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Amusement, -0.080 became -0.084.
Quality. Mean predicted overall quality across the segments went 2.80 → 3.13 (+0.33) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.413 → 0.496+0.083identity cos neighbours 0.413 → 0.496d_b rescored +0.206 → +0.581d_a rescored -0.080 → -0.084d_a mined -0.080d_b mined 0.206min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_ISSoAGrUrsItotal 27.2schain gain +2.9 dBseam step 0.4 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, average clarity
(amusement, teasing, pleasure ecstasy · normal-paced, energised, neutral tension, casual)(ahem) Uh, every time. They, they need a mechanic to get you from verse to verse. Right? And it almost, it's almost like, okay, who cares? Just, yeah, the next room is over here. You know, and you're just like, do you wanna just teleport to the next fight? Yeah! (chuckle) And that's how Bloody Palace (ahem) works! You just wanna run into the light and just get there? Yeah, okay, cool.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as amusement, teasing, pleasure ecstasy; style: casual, playful; average recording, some background noise; mildly explicit content; genuineness 6.0/6; vocal-burst blend 5.1/10; 21.7s, EN.
EN_ISSoAGrUrsI_W000099 · in -18.0 dBFS · gain -2.0 dB · emolia-00784
(disappointment, astonishment surprise, shame·measured, normally alert, slightly relaxed, casual)I did it too late. Now this is where they realize they've ziplined you too hard.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as disappointment, astonishment surprise, shame; style: casual, conversational; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 1.3/10; 5.7s, EN.
EN_ISSoAGrUrsI_W000100 · in -20.1 dBFS · gain +0.1 dB · emolia-00784
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Amusement drifts down from 0.99 to 0.86 (-0.13), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.11, then +0.10 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.66 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.66 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.66, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 52 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.642 before conversion and 0.784 after — it rose by 0.142. Neighbour-to-neighbour the worst pair went 0.642 → 0.784. (The earlier render, with segment 1 left raw, scores 0.481 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.195 in the original and +0.603 after conversion — 309 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Amusement, -0.123 became +0.018.
Quality. Mean predicted overall quality across the segments went 2.73 → 3.22 (+0.49) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.642 → 0.784+0.142identity cos neighbours 0.642 → 0.784d_b rescored +0.195 → +0.603d_a rescored -0.123 → +0.018d_a mined -0.126d_b mined 0.204min_cos_consec (site) 0.6647min_cos_anchor (site) 0.6647dataset podcastlang enspeaker 620540total 51.2schain gain +2.8 dBseam step 2.3 dBcrossfades 100/100 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an elderly masculine voice · slightly cool, below-average recording, moderately variable, slurred, audible breath
(amusement, teasing, contempt · fast, energised, tense, casual)On Ancestry.com claiming it's two percent of African that she got on her system. Like this gives her a plate at the goddamn cookout. You and your albino potato salad (breathy giggle) can go somewhere, girl.
full caption & clip details
An elderly masculine voice; delivery is energised, fast, tense, moderately variable; timbre is slightly cool, very dark, slightly rough, slightly thin; slurred, some disfluency, very wide pitch range, audible breath; affect is positive, slightly dominant, slightly guarded; reads as amusement, teasing, contempt; style: casual, storytelling; below-average recording, some background noise; genuineness 5.0/6; vocal-burst blend 8.3/10; 14.1s, EN.
620540_00039712 · in -47.2 dBFS · gain +27.2 dB · podcast-01165
(teasing, anger, impatience and irritability·slow, very low-energy, tense, storytelling)We don't want it, we don't need it, okay? You're not welcome to Wakanda. You can't come here. You cannot have a plate at the cookout at all. You cannot. If you came to my cookout, no. You cannot. Bring Chinese food. Cannot get a plate, because you are not (wistful sigh) welcome in
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, tense, moderately variable; timbre is slightly cool, dark, rough, thin; slurred, frequent disfluency, wide pitch range, audible breath; affect is negative, slightly dominant, fairly guarded; reads as teasing, anger, impatience and irritability; style: storytelling, cartoonish; below-average recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.2/10; 21.5s, EN.
620540_00041128 · in -48.8 dBFS · gain +28.8 dB · podcast-01162
(disappointment, impatience and irritability, bitterness·fast, very low-energy, neutral tension, casual)No, it was just us being happy to have to have something and for to have a budget and for it to look right and for it to be rare represented. We wouldn't it wasn't we weren't competing against we don't fuck Star Wars. We was just happy to be great for once. No,
full caption & clip details
A young adult masculine voice; delivery is very low-energy, fast, neutral tension, moderately variable; timbre is slightly cool, dark, rough, thin; slurred, some disfluency, wide pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, impatience and irritability, bitterness; style: casual, cartoonish; below-average recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.3/10; 15.8s, EN.
620540_00045048 · in -51.4 dBFS · gain +31.4 dB · podcast-01165
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Distress barely moves at all, sitting near 0.99 throughout.
It takes 2 clips to get there. Clip to clip the moves are +0.20 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 28 s · german · mls
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.921 before conversion and 0.905 after — it fell by 0.016. Neighbour-to-neighbour the worst pair went 0.921 → 0.905. (The earlier render, with segment 1 left raw, scores 0.815 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.204 in the original and +0.603 after conversion — 295 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Distress, -0.001 became +0.002.
Quality. Mean predicted overall quality across the segments went 3.04 → 3.27 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.921 → 0.905-0.016identity cos neighbours 0.921 → 0.905d_b rescored +0.204 → +0.603d_a rescored -0.001 → +0.002d_a mined -0.001d_b mined 0.204min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang germanspeaker 10148total 28.1schain gain +3.4 dBseam step 1.5 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, slightly relaxed, clear, light breath
(distress, sadness, malevolence malice · slow, very low-energy, fairly steady, ASMR)es werden zeiten des hungers und der armut kommen wie für das ganze volk so auch für jeden einzelnen menschen das ist klar
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as distress, sadness, malevolence malice; style: ASMR, monologue; very good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.1/10; 11.4s, GERMAN.
10148_11442_000813 · in -30.2 dBFS · gain +10.2 dB · mls-00008
(disappointment, bitterness, fear·normal-paced, normally alert, moderately variable, narration)sie mögen sagen was sie wollen der leib hängt doch nur von der seele ab wie kann man nur erwarten daß alles nach wunsch gehe denken sie nicht an die toten seelen sondern an ihre eigene lebendige seele und betreten sie mit gottes hilfe den neuen weg
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as disappointment, bitterness, fear; style: narration, storytelling; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 0.0/10; 16.9s, GERMAN.
10148_11442_000752 · in -27.8 dBFS · gain +7.8 dB · mls-00008
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.80, higher than 80 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Distress barely moves at all, sitting near 0.95 throughout.
It takes 2 clips to get there. Clip to clip the moves are +0.20 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 32 s · spanish · mls
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.968 before conversion and 0.942 after — it fell by 0.026. Neighbour-to-neighbour the worst pair went 0.968 → 0.942. (The earlier render, with segment 1 left raw, scores 0.785 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.200 in the original and +0.208 after conversion — 104 % of the delta retained, which is essentially all of it. On the other named axis, Distress, +0.036 became +0.059.
Quality. Mean predicted overall quality across the segments went 2.99 → 3.21 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.968 → 0.942-0.026identity cos neighbours 0.968 → 0.942d_b rescored +0.200 → +0.208d_a rescored +0.036 → +0.059d_a mined 0.036d_b mined 0.200min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang spanishspeaker 10246total 31.2schain gain +2.0 dBseam step 0.6 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normal-paced, normally alert, slightly relaxed
(distress, sadness, pain · no disfluency, monologue, narration)pero es justo añadir que habia olvidado en sus cálculos el reposo forzado de los domingos y dias de fiesta que en diez y nueve años hacian una disminucion de veinticuatro francos próximamente
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as distress, sadness, pain; style: monologue, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.6/10; 13.5s, SPANISH.
10246_11240_000374 · in -26.9 dBFS · gain +7.0 dB · mls-00033
(disappointment, emotional numbness, bitterness·almost no disfluency, narration, cartoonish)además esta masita habia sido reducida por varias retenciones á la suma de ciento nueve francos y quince sueldos que le habian sido entregados á su salida pero él no comprendia esto y se creia perjudicado digamos la palabra robado al
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, emotional numbness, bitterness; style: narration, cartoonish; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 4.0/10; 17.9s, SPANISH.
10246_11240_000532 · in -27.2 dBFS · gain +7.2 dB · mls-00033
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Teasing drifts down from 1.00 to 0.36 (-0.63), which the rule did not require.
It takes 2 clips to get there. Clip to clip the moves are +0.20 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 29 s · french · mls
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.931 before conversion and 0.926 after — it fell by 0.005. Neighbour-to-neighbour the worst pair went 0.931 → 0.926. (The earlier render, with segment 1 left raw, scores 0.831 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.201 in the original and +0.139 after conversion — 69 % of the delta retained. On the other named axis, Teasing, -0.634 became -0.636.
Quality. Mean predicted overall quality across the segments went 3.13 → 3.29 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.931 → 0.926-0.005identity cos neighbours 0.931 → 0.926d_b rescored +0.201 → +0.139d_a rescored -0.634 → -0.636d_a mined -0.634d_b mined 0.201min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang frenchspeaker 12541total 28.2schain gain +1.5 dBseam step 0.4 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, some disfluency
(teasing, shame, contentment · measured, didactic, monologue)où déjeunez vous harry chez tante agathe je me suis invité avec m gray c'est son dernier protégé bah dites donc à votre tante agathe harry de ne plus m'assommer avec ses oeuvres de charité
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as teasing, shame, contentment; style: didactic, monologue; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 2.6/10; 15.7s, FRENCH.
12541_12317_000180 · in -27.0 dBFS · gain +7.0 dB · mls-00048
(disappointment, distress, contempt·normal-paced, monologue, narration)j'en suis excédé la bonne femme croit-elle donc que je n'aie rien de mieux à faire que de signer des chèques en faveur de ses vilains drôles très bien oncle george je le lui dirai mais cela n'aura aucun effet
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, distress, contempt; style: monologue, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.4/10; 12.8s, FRENCH.
12541_12317_000334 · in -24.1 dBFS · gain +4.1 dB · mls-00048
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Teasing drifts down from 1.00 to 0.94 (-0.06), which the rule did not require.
It takes 2 clips to get there. Clip to clip the moves are +0.20 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 36 s · italian · mls
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.946 before conversion and 0.909 after — it fell by 0.038. Neighbour-to-neighbour the worst pair went 0.946 → 0.909. (The earlier render, with segment 1 left raw, scores 0.807 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.202 in the original and +0.063 after conversion — 31 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Teasing, -0.055 became -0.021.
Quality. Mean predicted overall quality across the segments went 3.03 → 3.23 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.946 → 0.909-0.038identity cos neighbours 0.946 → 0.909d_b rescored +0.202 → +0.063d_a rescored -0.055 → -0.021d_a mined -0.055d_b mined 0.202min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang italianspeaker 4971total 35.9schain gain +3.3 dBseam step 1.1 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(teasing, contempt, anger · moderately variable, wide pitch range, cartoonish, playful)e difatti eleonora s'era provata a intercedere ma il mezzadro ah nonononò ossequio rispetto tutto il rispetto per la signorina ma anche preghiera di non immischiarsi ed allora essa un po per pietà un po per ridere un po per darsi da fare s'era messa ad ajutare quel povero giovanotto fin dove poteva
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as teasing, contempt, anger; style: cartoonish, playful; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 5.3/10; 19.8s, ITALIAN.
4971_4125_000199 · in -20.7 dBFS · gain +0.7 dB · mls-00072
(disappointment, shame, contempt ·fairly steady, moderate pitch range, monologue, narration)lo faceva ogni dopo pranzo venir su coi libri e i quaderni della scuola egli saliva impacciato e vergognoso perché s'accorgeva che la padrona prendeva a goderselo per la sua balordaggine per la sua durezza di mente ma che poteva farci il padre voleva così
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, shame, contempt; style: monologue, narration; good recording, quiet background; genuineness 2.0/6; vocal-burst blend 3.4/10; 16.4s, ITALIAN.
4971_4125_000163 · in -22.6 dBFS · gain +2.6 dB · mls-00072
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Elation drifts down from 0.98 to 0.92 (-0.07), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.00, then -0.00, then +0.00 — not a clean run: step 3 moves back the other way by 0.00 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.91 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.91 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 129 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.907 before conversion and 0.894 after — it fell by 0.013. Neighbour-to-neighbour the worst pair went 0.874 → 0.919. (The earlier render, with segment 1 left raw, scores 0.790 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.201 in the original and +0.599 after conversion — 298 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Elation, -0.067 became -0.044.
Quality. Mean predicted overall quality across the segments went 2.69 → 3.10 (+0.40) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.907 → 0.894-0.013identity cos neighbours 0.874 → 0.919d_b rescored +0.201 → +0.599d_a rescored -0.067 → -0.044d_a mined -0.067d_b mined 0.201min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_ykDYkgzsXE0total 128.0schain gain +3.1 dBseam step 0.9 dBcrossfades 100/100/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · slightly cool, energised, moderately variable, some disfluency
(elation, hope enthusiasm optimism, interest · brisk, neutral tension, average clarity, casual)To finish off the squad up top, we have Neymar, Baal and Benzema. Now these three are very good players. Benzema's currently going for 21k coins at the moment but once the crash hits,
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as elation, hope enthusiasm optimism, interest; style: casual, dramatic; below-average recording, quiet background; mildly explicit content; genuineness 3.6/6; vocal-burst blend 10.0/10; 23.6s, EN.
EN_ykDYkgzsXE0_W000006 · in -12.3 dBFS · gain -7.7 dB · emolia-01637
(disappointment, hope enthusiasm optimism, bitterness·fast, neutral tension, average clarity, casual)He'll probably drop to anywhere between 10 to 15k coins. And since he's a good player, he'll definitely be wanted by. And Bear was currently going for 131k coins. And because of people buying packs, wanting to get coins quick, they'll probably start selling him like crazy. And he might drop to even 100k coins, or even like 90k coins. So yea, he'll definitely be good for this squad. And finally, we have Neymar. Neymar's normal card is currently priced at 160k coins. Bear in mind, he's got 3 informs on top of that. And he will be getting a team of the year. So this card,
full caption & clip details
A young adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as disappointment, hope enthusiasm optimism, bitterness; style: casual, dramatic; below-average recording, quiet background; mildly explicit content; genuineness 2.9/6; vocal-burst blend 10.0/10; 29.7s, EN.
EN_ykDYkgzsXE0_W000007 · in -12.1 dBFS · gain -7.9 dB · emolia-01637
(triumph, disappointment, pleasure ecstasy·brisk, neutral tension, average clarity, casual)But after the crash, after his team of the year has been released, this guy will probably drop by 10k coins. He might end up even going for around about 20k coins, which is why he's perfect for this side. And Tiago Silva, currently going for 24k coins. Not the most I know, but he'll probably drop by a good 8k coins. So yea, it'll be another card to get. Overall, the team is pretty strong. Got a few good players in there. Definitely players that a lot of you guys probably want to use. Once the market crash hit, this is the team you want to build, guys. I mean, you'll save money. And once team of the year is over, you'll be able to make yourself some
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as triumph, disappointment, pleasure ecstasy; style: casual, dramatic; average recording, quiet background; mildly explicit content; genuineness 2.6/6; vocal-burst blend 9.6/10; 30.0s, EN.
EN_ykDYkgzsXE0_W000009 · in -11.2 dBFS · gain -8.8 dB · emolia-01637
(pride, triumph, elation·fast, slightly relaxed, average clarity, dramatic)I know for most people it's still kind of expensive, but I know you guys are going to be opening packs, you guys are going to get some insane players. So if you do and you get those coins from your players, make sure to put 300k aside and build this team because it's a really strong team and it will also make you a bit of profit once team of the year is over.
full caption & clip details
A young adult masculine voice; delivery is energised, fast, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as pride, triumph, elation; style: dramatic, casual; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 8.0/10; 30.0s, EN.
EN_ykDYkgzsXE0_W000010 · in -10.1 dBFS · gain -9.9 dB · emolia-01637
(disappointment, bitterness, anger· fast, neutral tension, somewhat unclear, casual)Well, I say little bench, but look at the players that are on here with players that will drop in price as always. Lacazette, he's surely going to drop. Messi always drops every year with team of the year. Same with Ronaldo. They always get team of the year. So yea, they will both drop in price. Kosta will drop in price.
full caption & clip details
A young adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, slightly bright, very rough, thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is elated, slightly dominant, slightly guarded; reads as disappointment, bitterness, anger; style: casual, dramatic; below-average recording, some background noise; mildly explicit content; genuineness 5.6/6; vocal-burst blend 10.0/10; 15.4s, EN.
EN_ykDYkgzsXE0_W000011 · in -9.9 dBFS · gain -10.1 dB · emolia-01637
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.80, higher than 80 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Elation barely moves at all, sitting near 0.99 throughout.
It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.04 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.89 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.89 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 56 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.867 before conversion and 0.851 after — it fell by 0.016. Neighbour-to-neighbour the worst pair went 0.866 → 0.817. (The earlier render, with segment 1 left raw, scores 0.670 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.203 in the original and +0.603 after conversion — 296 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Elation, +0.009 became +0.019.
Quality. Mean predicted overall quality across the segments went 2.87 → 3.11 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.867 → 0.851-0.016identity cos neighbours 0.866 → 0.817d_b rescored +0.203 → +0.603d_a rescored +0.009 → +0.019d_a mined 0.009d_b mined 0.204min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00049_S00296total 55.4schain gain +3.3 dBseam step 1.7 dBcrossfades 100/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · slightly cool, slightly bright, below-average recording, noisy background, brisk, slightly tense, moderately variable, normal breath
(elation, interest, triumph · highly aroused, almost no disfluency, clear, dramatic)What can RNG take? They've already opened up the base at 25 minutes and VDD doing what he can to push in the top lane. Chips away 60, 70% of that tower before he backs away, but the game is blown wide open by RNG.
full caption & clip details
A young adult masculine voice; delivery is highly aroused, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, almost no disfluency, very wide pitch range, normal breath; affect is elated, very dominant, fairly guarded; reads as elation, interest, triumph; style: dramatic, ranting; below-average recording, noisy background; genuineness 1.9/6; vocal-burst blend 3.7/10; 15.4s, EN.
EN_B00049_S00296_W000020 · in -22.8 dBFS · gain +2.8 dB · emolia-01198
(pain, triumph, astonishment surprise·energised, some disfluency, clear, ranting)My eyes are on Khan. Yesterday, in a game-defining match against Flash Wolves, Khan found Flash Wolves, AD Carry, Betty twice with a Flash-Rupture-Feast combo and just sealed the deal and sent them to the finals. That'll be much more difficult against a Black Shield Ezreal. However, Khan is the man that is now being jumped on.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, slightly rough, balanced body; clear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, fairly guarded; reads as pain, triumph, astonishment surprise; style: ranting, dramatic; below-average recording, noisy background; mildly explicit content; genuineness 2.8/6; vocal-burst blend 4.3/10; 21.3s, EN.
EN_B00049_S00296_W000021 · in -24.8 dBFS · gain +4.8 dB · emolia-01198
(elation, disappointment, triumph ·highly aroused, some disfluency, average clarity, casual)I'm not gonna lie, there were moments around the 20 to 25 minute mark where I thought RNG had this. I thought they were just gonna clean these team fights up and Kingzone just shut me the hell up in their own jungle.
full caption & clip details
A young adult masculine voice; delivery is highly aroused, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, slightly guarded; reads as elation, disappointment, triumph; style: casual, ranting; below-average recording, noisy background; mildly explicit content; genuineness 3.2/6; vocal-burst blend 4.0/10; 19.0s, EN.
EN_B00049_S00296_W000022 · in -22.0 dBFS · gain +2.0 dB · emolia-01198
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Emotional Numbness drifts down from 0.99 to 0.82 (-0.17), which the rule did not require.
It takes 2 clips to get there. Clip to clip the moves are +0.20 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 34 s · spanish · mls
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.967 before conversion and 0.957 after — it fell by 0.010. Neighbour-to-neighbour the worst pair went 0.967 → 0.957. (The earlier render, with segment 1 left raw, scores 0.858 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.202 in the original and +0.602 after conversion — 298 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, -0.172 became -0.162.
Quality. Mean predicted overall quality across the segments went 3.20 → 3.30 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.967 → 0.957-0.010identity cos neighbours 0.967 → 0.957d_b rescored +0.202 → +0.602d_a rescored -0.172 → -0.162d_a mined -0.172d_b mined 0.202min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang spanishspeaker 3946total 33.2schain gain -1.7 dBseam step 0.2 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normal-paced, normally alert, slightly relaxed
(emotional numbness, malevolence malice, bitterness · steady, fairly narrow pitch, monologue, narration)y que sobre ello pusiese grandes penas e vino á cuba el mismo oidor y hizo sus diligencias y protestaciones como le era mandado por la real audiencia para que no saliese con su intencion el velazquez
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, malevolence malice, bitterness; style: monologue, narration; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 1.0/10; 14.2s, SPANISH.
3946_11219_006661 · in -25.2 dBFS · gain +5.2 dB · mls-00032
(disappointment, anger, bitterness ·fairly steady, moderate pitch range, monologue, narration)y por mas penas y requirimien tos que le hizo é puso no aprovechó cosa ninguna porque como el diego velazquez era tan favorecido del obispo de burgos y habla gastado quanto tenia en hacer aquella gente de guerra contra nosotros no tuvo todos aquellos requirimientos que hicieron en una castañeta
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, anger, bitterness; style: monologue, narration; good recording, quiet background; genuineness 0.2/6; vocal-burst blend 3.9/10; 19.2s, SPANISH.
3946_11219_006557 · in -25.1 dBFS · gain +5.1 dB · mls-00032
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.80, higher than 80 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Emotional Numbness barely moves at all, sitting near 0.99 throughout.
It takes 2 clips to get there. Clip to clip the moves are +0.20 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 30 s · spanish · mls
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.916 before conversion and 0.897 after — it fell by 0.019. Neighbour-to-neighbour the worst pair went 0.916 → 0.897. (The earlier render, with segment 1 left raw, scores 0.768 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.201 in the original and +0.115 after conversion — 57 % of the delta retained. On the other named axis, Emotional Numbness, -0.012 became +0.001.
Quality. Mean predicted overall quality across the segments went 3.09 → 3.36 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.916 → 0.897-0.019identity cos neighbours 0.916 → 0.897d_b rescored +0.201 → +0.115d_a rescored -0.012 → +0.001d_a mined -0.012d_b mined 0.202min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang spanishspeaker 3946total 29.7schain gain -1.2 dBseam step 0.5 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, slightly relaxed, steady, almost no disfluency, clear
(emotional numbness, triumph, pain · measured, subdued, fairly narrow pitch, monologue)con que enriquecieron su nave pareciéndoles que en la hermosura de auristela llevaban un precioso y nunca visto rescate quise llegar con mi barca á hablar con el capitan de los vencedores pero como mi ventura andaba siempre en los aires uno de tierra sopló y hizo apartar el navío
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, triumph, pain; style: monologue, narration; average recording, quiet background; genuineness 0.4/6; vocal-burst blend 2.6/10; 19.7s, SPANISH.
3946_10158_002847 · in -21.5 dBFS · gain +1.5 dB · mls-00024
(disappointment, helplessness, sourness·normal-paced, normally alert, moderate pitch range, monologue)no pude llegar á él ni ofrecer imposibles por el rescate de la presa y así fué forzoso el volvernos sin ninguna esperanza de cobrar nuestra pérdida
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, helplessness, sourness; style: monologue, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 10.2s, SPANISH.
3946_10158_001696 · in -22.6 dBFS · gain +2.6 dB · mls-00024
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.21.
Nothing was asked of the other axis, and in fact Pain barely moves at all, sitting near 1.00 throughout.
It takes 2 clips to get there. Clip to clip the moves are +0.21 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 34 s · italian · mls
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.665 before conversion and 0.785 after — it rose by 0.120. Neighbour-to-neighbour the worst pair went 0.665 → 0.785. (The earlier render, with segment 1 left raw, scores 0.660 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.206 in the original and +0.186 after conversion — 90 % of the delta retained, which is essentially all of it. On the other named axis, Pain, -0.012 became -0.014.
Quality. Mean predicted overall quality across the segments went 3.09 → 3.35 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.665 → 0.785+0.120identity cos neighbours 0.665 → 0.785d_b rescored +0.206 → +0.186d_a rescored -0.012 → -0.014d_a mined -0.012d_b mined 0.206min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang italianspeaker 1595total 33.6schain gain +1.2 dBseam step 0.8 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-bright, almost no disfluency, clear
(pain, shame, triumph · measured, normally alert, slightly relaxed, narration)di questa lingua bem questa cosa di numeri come si stia e se così la prosa come il verso toscano
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, fairly guarded; reads as pain, shame, triumph; style: narration, storytelling; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.0/10; 15.9s, ITALIAN.
1595_5071_000245 · in -25.9 dBFS · gain +5.8 dB · mls-00073
(disappointment, shame, bitterness·fast, energised, neutral tension, cartoonish)la sua parte e in che modo la si abbia per essere assai facile da vedere ma lontana dal nostro proponimento ora con esso voi non intendo di disputarla anzi confessando quello esser vero che ne diceste non tanto perché sia vero quanto perché si veda ciò che ne segue
full caption & clip details
An elderly masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, rough, thin; clear, almost no disfluency, very wide pitch range, audible breath; affect is positive, slightly dominant, fairly guarded; reads as disappointment, shame, bitterness; style: cartoonish, storytelling; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 6.0/10; 17.9s, ITALIAN.
1595_5071_000106 · in -21.8 dBFS · gain +1.8 dB · mls-00073
This chain comes from the one-sided rule: only Disappointment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Disappointment strongly present — 0.80, higher than 80 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Pain drifts down from 0.99 to 0.86 (-0.13), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.16, then -0.01, then +0.05, then +0.00 — not a clean run: step 2 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 71 s · spanish · mls
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.938 before conversion and 0.941 after — it rose by 0.003. Neighbour-to-neighbour the worst pair went 0.904 → 0.919. (The earlier render, with segment 1 left raw, scores 0.840 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.201 in the original and +0.169 after conversion — 84 % of the delta retained, which is most of it. On the other named axis, Pain, -0.130 became -0.251.
Quality. Mean predicted overall quality across the segments went 2.94 → 3.22 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.938 → 0.941+0.003identity cos neighbours 0.904 → 0.919d_b rescored +0.201 → +0.169d_a rescored -0.130 → -0.251d_a mined -0.130d_b mined 0.201min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang spanishspeaker 10246total 69.5schain gain +2.5 dBseam step 0.7 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-bright, normally alert, slightly relaxed, fairly steady, clear, moderate pitch range, light breath
(pain, malevolence malice, bitterness · normal-paced, no disfluency, narration, monologue)aquella carne abandonada á los reptiles del lago era carne suya aquella envoltura de materia vivero da ianguijuelas y gusanos era el fruto de sus arrebatos apasionados de bu amor insaciable en el silencio de ia noche
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain, malevolence malice, bitterness; style: narration, monologue; good recording, quiet background; genuineness 0.8/6; vocal-burst blend 2.8/10; 15.3s, SPANISH.
10246_11643_001442 · in -25.8 dBFS · gain +5.8 dB · mls-00036
(distress, helplessness, emotional numbness·measured, no disfluency, narration, monologue)la enormidad del crimen le abrumaba nada de excusas no debía buscar pretextos como otras veces para seguir adelante era un miserable indigno de vivir
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as distress, helplessness, emotional numbness; style: narration, monologue; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 2.2/10; 11.9s, SPANISH.
10246_11643_001337 · in -26.4 dBFS · gain +6.4 dB · mls-00036
(sadness, emotional numbness, sourness· measured, no disfluency, narration, cartoonish)una rama eeca del árbol de los palomas siempre recto siempre vigoroso coo aspereza salvaje pero sano en medio de su aislamiento la mala rama debía desaparecer
full caption & clip details
An elderly feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sadness, emotional numbness, sourness; style: narration, cartoonish; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 4.1/10; 12.6s, SPANISH.
10246_11643_000200 · in -26.0 dBFS · gain +6.0 dB · mls-00036
(shame, anger, disappointment· measured, almost no disfluency, narration, cartoonish)su abuelo tenía razón al despreciarlo su padre su pobre padre al que ahora contemplaba con la grandeza de ios santos hacía bien en repelerle como un brote infame de su existencia la infeliz borda con su vergonzoso origen era más hija de los palomas que é
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as shame, anger, disappointment; style: narration, cartoonish; average recording, no background noise; genuineness 1.4/6; vocal-burst blend 5.1/10; 19.2s, SPANISH.
10246_11643_000554 · in -26.2 dBFS · gain +6.2 dB · mls-00036
(disappointment, sadness, helplessness· measured, no disfluency, storytelling, narration)qué había hecho durante su vida nada su voluntad sólo tenía fuerzas para huir del trabajo ei desdichado sangonera había sido mejor que él
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as disappointment, sadness, helplessness; style: storytelling, narration; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.4/10; 11.3s, SPANISH.
10246_11643_000689 · in -26.5 dBFS · gain +6.5 dB · mls-00036