This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_c-mls-AB2.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words
are a real non-speech sound, printed where it happens.
The full generated caption for any clip is under “full caption & clip
details”. Its perceived-gender and background-noise clauses were re-rendered from
the numeric buckets, because the versions stored in the corpus index had those two
ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
The tags and captions describe the ORIGINAL clips — they are the
corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the
converted audio are the ones printed in each card's own paragraph and chips.
This chain comes from the two-sided rule: it only counts if both emotions move — Contentment down and Concentration up — by at least 0.25 each.
The chain starts with Concentration clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.31.
At the same time Contentment goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.56 (higher than 56 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.07, then +0.13, then +0.12 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 67 s · polish · mls
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.943 before conversion and 0.954 after — it rose by 0.011. Neighbour-to-neighbour the worst pair went 0.962 → 0.961. (The earlier render, with segment 1 left raw, scores 0.817 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.312 in the original and +0.339 after conversion — 109 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contentment, -0.363 became -0.227.
Quality. Mean predicted overall quality across the segments went 3.13 → 3.47 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.943 → 0.954+0.011identity cos neighbours 0.962 → 0.961d_b rescored +0.312 → +0.339d_a rescored -0.363 → -0.227d_a mined -0.362d_b mined 0.311min_cos_consec (site) 0.9635min_cos_anchor (site) 0.9448dataset mlslang polishspeaker 6892total 65.8schain gain +0.5 dBseam step 0.9 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a child masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, quiet background, normal-paced, normally alert
(contentment, triumph · some disfluency, monologue, narration)wąska kanapka obita skórą dwa krzesła również skórą obite duża blaszana miednica i mała szafa ciemnowiśniowej barwy stanowiły umeblowanie pokoju który ze względu na swoję długość i mrok w nim panujący zdawał się być podobniejszym do grobu aniżeli do mieszkania
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contentment, triumph; style: monologue, narration; good recording, quiet background; genuineness 0.3/6; vocal-burst blend 3.1/10; 18.4s, POLISH.
6892_8764_000241 · in -26.2 dBFS · gain +6.2 dB · mls-00111
(pride, malevolence malice, shame· some disfluency, narration, monologue)równie jak pokój nie zmieniły się od ćwierć wieku zwyczaje pana ignacego rano budził się zawsze o szóstej przez chwilę słuchał czy idzie leżący na krześle zegarek i spoglądał na skazówki które tworzyły jednę linią prostą
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, malevolence malice, shame; style: narration, monologue; good recording, quiet background; genuineness 0.7/6; vocal-burst blend 1.1/10; 14.8s, POLISH.
6892_8764_000148 · in -26.7 dBFS · gain +6.7 dB · mls-00111
(triumph, relief·almost no disfluency, narration, monologue)chciał wstać spokojnie bez awantur ale że chłodne nogi i nieco zesztywniałe ręce nie okazywały się dość uległemi jego woli więc zrywał się nagle wyskakiwał na środek pokoju i rzuciwszy na łóżko szlafmycę biegł pod piec do wielkiej miednicy
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, relief; style: narration, monologue; good recording, quiet background; genuineness 0.3/6; vocal-burst blend 2.7/10; 15.6s, POLISH.
6892_8764_001691 · in -26.3 dBFS · gain +6.3 dB · mls-00111
(concentration, interest·some disfluency, narration, monologue)w której mył się od stóp do głów rżąc i parskając jak wiekowy rumak szlachetnej krwi któremu przypomniał się wyścig podczas obrządku wycierania się kosmatemi ręcznikami z upodobaniem patrzał na swoje chude łydki i zarośnięte piersi mrucząc no przecie nabieram ciała
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, interest; style: narration, monologue; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 2.7/10; 17.7s, POLISH.
6892_8764_000379 · in -26.6 dBFS · gain +6.6 dB · mls-00111
This chain comes from the two-sided rule: it only counts if both emotions move — Jealousy and Envy down and Emotional Numbness up — by at least 0.25 each.
The chain starts with Emotional Numbness clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.25.
At the same time Jealousy and Envy goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.44. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.09, then +0.18, then -0.04, then +0.03 — not a clean run: step 3 moves back the other way by 0.04 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 76 s · spanish · mls
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.926 before conversion and 0.934 after — it rose by 0.008. Neighbour-to-neighbour the worst pair went 0.927 → 0.917. (The earlier render, with segment 1 left raw, scores 0.821 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.254 in the original and +0.218 after conversion — 86 % of the delta retained, which is most of it. On the other named axis, Jealousy and Envy, -0.439 became -0.492.
Quality. Mean predicted overall quality across the segments went 3.06 → 3.26 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.926 → 0.934+0.008identity cos neighbours 0.927 → 0.917d_b rescored +0.254 → +0.218d_a rescored -0.439 → -0.492d_a mined -0.439d_b mined 0.255min_cos_consec (site) 0.9389min_cos_anchor (site) 0.9238dataset mlslang spanishspeaker 12367total 74.7schain gain +2.3 dBseam step 0.8 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, balanced body, measured, slightly relaxed, steady, fairly narrow pitch
(jealousy and envy, contentment, disgust · subdued, no disfluency, slurred, narration)quisiera que hubieras estado allí para observar la maniobra y eso de decir que al capitán le gustan igualmente enriqueta y luisa es tonto porque le gusta sin duda enriqueta mucho más
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, contentment, disgust; style: narration, monologue; good recording, quiet background; genuineness 0.1/6; vocal-burst blend 0.4/10; 13.3s, SPANISH.
12367_11991_000132 · in -26.4 dBFS · gain +6.3 dB · mls-00039
(disappointment, bitterness, anger· subdued, some disfluency, somewhat unclear, monologue)pero este carlos es tan prosaico me hubiera gustado que estuvieras con nosotras ayer tarde para haber juzgado por ti misma y estoy segura de que pensarías como yo a menos de que quisieras llevarme la contraria conque ana asistiera a la comida en casa de
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as disappointment, bitterness, anger; style: monologue, whispered; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 5.5/10; 19.2s, SPANISH.
12367_11991_000126 · in -26.3 dBFS · gain +6.3 dB · mls-00039
(emotional numbness, fatigue exhaustion, contentment·normally alert, no disfluency, slurred, monologue)le hubiera bastado para observar todas estas cosas pero se quedó en casa con un pretexto en el que se complicaban un dolor de cabeza y un retroceso en la enfermedad de carlitos
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, no disfluency, fairly narrow pitch, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, fatigue exhaustion, contentment; style: monologue, whispered; average recording, no background noise; genuineness 0.1/6; vocal-burst blend 1.7/10; 13.6s, SPANISH.
12367_11991_000080 · in -26.4 dBFS · gain +6.4 dB · mls-00039
(emotional numbness, triumph, contentment ·subdued, no disfluency, somewhat unclear, narration)era su único propósito evitar al capitán wentwortfa pero el resultado había sido eludir el ser tomada como árbitro al mismo tiempo que disfrutar de una tarde tranquila
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, triumph, contentment; style: narration, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.8/10; 13.1s, SPANISH.
12367_11991_000077 · in -26.8 dBFS · gain +6.8 dB · mls-00039
(emotional numbness, malevolence malice, thankfulness gratitude· subdued, almost no disfluency, somewhat unclear, monologue)respecto a los designios del capitán wentworth pensaba ella que debía importar más a federico resolver pronto su preferencia si no quería comprometer ja felicidad de las dos muchachas o desdorar su propio honor
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, malevolence malice, thankfulness gratitude; style: monologue, narration; average recording, quiet background; genuineness 0.1/6; vocal-burst blend 1.7/10; 16.3s, SPANISH.
12367_11991_000025 · in -26.5 dBFS · gain +6.5 dB · mls-00039
This chain comes from the two-sided rule: it only counts if both emotions move — Pride down and Contempt up — by at least 0.25 each.
The chain starts with Contempt clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.27.
At the same time Pride goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.12 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 44 s · spanish · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.938 before conversion and 0.904 after — it fell by 0.034. Neighbour-to-neighbour the worst pair went 0.913 → 0.874. (The earlier render, with segment 1 left raw, scores 0.795 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contempt moved +0.266 in the original and +0.457 after conversion — 172 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Pride, -0.309 became -0.296.
Quality. Mean predicted overall quality across the segments went 3.00 → 3.20 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.938 → 0.904-0.034identity cos neighbours 0.913 → 0.874d_b rescored +0.266 → +0.457d_a rescored -0.309 → -0.296d_a mined -0.309d_b mined 0.265min_cos_consec (site) 0.9375min_cos_anchor (site) 0.9423dataset mlslang spanishspeaker 10246total 43.2schain gain +2.1 dBseam step 1.3 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · neutral-toned, neutral-bright, fairly smooth, normally alert, slightly relaxed, fairly steady, clear, moderate pitch range
(pride · measured, no disfluency, narration, monologue)en este punto el maestro de escuela impugnó igualmente el sermón y defendió con más calor ahinco y acierto á juanita es decía una muchacha discreta honrada y trabajadora
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride; style: narration, monologue; good recording, quiet background; genuineness 0.8/6; vocal-burst blend 1.8/10; 15.5s, SPANISH.
10246_11832_000209 · in -26.6 dBFS · gain +6.7 dB · mls-00038
(anger, sourness, shame·normal-paced, almost no disfluency, narration, cartoonish)dios la ha hecho hermosísima y casi estoy por decir que no sólo tiene derecho sino que tiene el deber de acicalarse y de realzar y mostrar la hermosura que dios le ha dado lo contrario sería ingratitud para con dios y desdeñar lo que enseña la parábola de los cinco talentos
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger, sourness, shame; style: narration, cartoonish; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 5.0/10; 16.7s, SPANISH.
10246_11832_001044 · in -25.0 dBFS · gain +5.0 dB · mls-00038
(contempt, astonishment surprise, malevolence malice· normal-paced, almost no disfluency, authoritative, monologue)talentos y extraño mucho que ustedes que han estado conmigo de fendiendo la propiedad individual se vuelvan aho ra contra mí y se pongan del lado de don policarpo para impugnar dicha propiedad
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, astonishment surprise, malevolence malice; style: authoritative, monologue; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.6/10; 11.3s, SPANISH.
10246_11832_000488 · in -24.4 dBFS · gain +4.3 dB · mls-00038
This chain comes from the two-sided rule: it only counts if both emotions move — Contempt down and Emotional Numbness up — by at least 0.25 each.
The chain starts with Emotional Numbness clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.28.
At the same time Contempt goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.50 (right about the corpus median), a change of -0.48. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.20, then -0.10, then +0.17 — not a clean run: step 2 moves back the other way by 0.10 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 62 s · french · mls
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.939 before conversion and 0.872 after — it fell by 0.067. Neighbour-to-neighbour the worst pair went 0.940 → 0.872. (The earlier render, with segment 1 left raw, scores 0.884 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.279 in the original and +0.345 after conversion — 124 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contempt, -0.476 became -0.064.
Quality. Mean predicted overall quality across the segments went 3.18 → 3.29 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.939 → 0.872-0.067identity cos neighbours 0.940 → 0.872d_b rescored +0.279 → +0.345d_a rescored -0.476 → -0.064d_a mined -0.476d_b mined 0.279min_cos_consec (site) 0.9598min_cos_anchor (site) 0.9495dataset mlslang frenchspeaker 12709total 61.2schain gain +0.3 dBseam step 1.3 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(contempt, sourness, awe · normal-paced, moderate pitch range, monologue, narration)om carlos est aussi un ouvrage de la jeunesse de schiller et cependant on le considère comme une composition du premier rang ce sujet de don carlos est un des plus dramatiques que l'histoire puisse offrir
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, sourness, awe; style: monologue, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.1/10; 13.0s, FRENCH.
12709_13781_004667 · in -26.4 dBFS · gain +6.4 dB · mls-00055
(sourness, awe, malevolence malice· normal-paced, fairly narrow pitch, monologue, narration)une jeune princesse fille de henri ï quitte la france et la cour brillante et chevaleresque du roi son père pour s'unir à un vieux tyran tellement sombre et sévère que le caractère même des espagnols fut altéré par son règne et que pendant longtemps la nation porta l'empreinte de son maître
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, awe, malevolence malice; style: monologue, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.2/10; 18.6s, FRENCH.
12709_13781_004091 · in -26.2 dBFS · gain +6.2 dB · mls-00055
(pain, sadness, fear·measured, fairly narrow pitch, monologue, narration)carlos fiancé d'abord à élisabeth l'aime encore quoiqu'elle soit devenue sa belle-mère la réformation et la révolte des paysbas ces grands événements politiques se mêlent à la catastrophe tragique de la condamnation du fils par le père
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain, sadness, fear; style: monologue, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.1/10; 16.0s, FRENCH.
12709_13781_005073 · in -27.4 dBFS · gain +7.4 dB · mls-00055
(emotional numbness· measured, fairly narrow pitch, monologue, newsreading)l'intérêt individuel et l'intérêt pumic se trouvent réunis au plus haut degré dans cette tragédie plusieurs écrivains ont traité ce sujet en france mais on n'a pu dans l'ancien régime le mettre sur le théâtre
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: monologue, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 14.3s, FRENCH.
12709_13781_004997 · in -27.0 dBFS · gain +7.0 dB · mls-00055
This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Shame up — by at least 0.25 each.
The chain starts with Shame clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.34.
At the same time Concentration goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.67 (higher than 67 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.19, then +0.00, then +0.15 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 55 s · dutch · mls
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.822 before conversion and 0.867 after — it rose by 0.045. Neighbour-to-neighbour the worst pair went 0.791 → 0.867. (The earlier render, with segment 1 left raw, scores 0.599 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.335 in the original and +0.246 after conversion — 73 % of the delta retained, which is most of it. On the other named axis, Concentration, -0.302 became -0.252.
Quality. Mean predicted overall quality across the segments went 2.96 → 3.38 (+0.42) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.822 → 0.867+0.045identity cos neighbours 0.791 → 0.867d_b rescored +0.335 → +0.246d_a rescored -0.302 → -0.252d_a mined -0.301d_b mined 0.335min_cos_consec (site) 0.8183min_cos_anchor (site) 0.8183dataset mlslang dutchspeaker 2450total 53.9schain gain +1.4 dBseam step 0.6 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an elderly masculine voice · neutral-toned, slightly dark, average recording, quiet background, measured, normally alert, slightly relaxed, fairly steady
(concentration, impatience and irritability · frequent disfluency, wide pitch range, audible breath, cartoonish)neen zeide mevrouw aarssen dat ware niet raadzaam geweest want zoo als de dafnis zegt in de minneklacht hier zeit u de ronde sta daer de schildwacht
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, thin; slurred, frequent disfluency, wide pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, impatience and irritability; style: cartoonish, storytelling; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 0.7/10; 15.0s, DUTCH.
2450_10026_003886 · in -29.1 dBFS · gain +9.1 dB · mls-00085
(malevolence malice, emotional numbness, concentration ·some disfluency, fairly narrow pitch, audible breath, narration)en zoo vervolgde sylvius bleef ik te scheveningen en zond u heden middag onzen getrouwen maarten ten einde te vernemen of de zaak in orde ware
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, thin; slurred, some disfluency, fairly narrow pitch, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as malevolence malice, emotional numbness, concentration; style: narration, monologue; average recording, quiet background; genuineness 0.9/6; vocal-burst blend 0.4/10; 11.0s, DUTCH.
2450_10026_002890 · in -28.4 dBFS · gain +8.4 dB · mls-00084
(malevolence malice, emotional numbness, concentration · some disfluency, fairly narrow pitch, audible breath, narration)en zoo vervolgde sylvius bleef ik te scheveningen en zond u heden middag onzen getrouwen maarten ten einde te vernemen of de zaak in orde ware
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, thin; slurred, some disfluency, fairly narrow pitch, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as malevolence malice, emotional numbness, concentration; style: narration, monologue; average recording, quiet background; genuineness 0.9/6; vocal-burst blend 0.4/10; 11.0s, DUTCH.
2450_10026_002890 · in -28.4 dBFS · gain +8.4 dB · mls-00084
(shame, jealousy and envy, disappointment·frequent disfluency, wide pitch range, normal breath, cartoonish)en zijt gy alleen gekomen vroeg buat en hoe hebt gy het op reis gehad vroeg zijn vrouw veroorlooft my een voor een uwe vragen te beantwoorden zeide sylvius mevrouw doet my de eer aan my te vragen hoe ik het gehad heb
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, wide pitch range, normal breath; affect is neutral, neutral stance, slightly guarded; reads as shame, jealousy and envy, disappointment; style: cartoonish, didactic; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.8/10; 17.5s, DUTCH.
2450_10026_000162 · in -28.9 dBFS · gain +8.9 dB · mls-00084
This chain comes from the two-sided rule: it only counts if both emotions move — Disgust down and Contemplation up — by at least 0.25 each.
The chain starts with Contemplation clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.36.
At the same time Disgust goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.13 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 55 s · french · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.945 before conversion and 0.913 after — it fell by 0.031. Neighbour-to-neighbour the worst pair went 0.924 → 0.888. (The earlier render, with segment 1 left raw, scores 0.830 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.359 in the original and +0.150 after conversion — 42 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Disgust, -0.330 became -0.643.
Quality. Mean predicted overall quality across the segments went 3.18 → 3.35 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.945 → 0.913-0.031identity cos neighbours 0.924 → 0.888d_b rescored +0.359 → +0.150d_a rescored -0.330 → -0.643d_a mined -0.330d_b mined 0.359min_cos_consec (site) 0.9275min_cos_anchor (site) 0.9491dataset mlslang frenchspeaker 12709total 54.7schain gain +1.3 dBseam step 1.1 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, measured, normally alert, slightly relaxed
(disgust, sadness, shame · formal, monologue)il se fût agi d'aller à la messe de paroisse que leur mise n'eût été pas plus décente seuls les gros cache-nez de couleurs vives noués sur la gorge et les vestes de bure jetées sur l'épaule en guise de pardessus avertissaient du voisinage des pays arctiques quand tu voudras jean rené prononça le capitaine
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, sadness, shame; style: formal, monologue; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.6/10; 19.8s, FRENCH.
12709_14178_000618 · in -29.5 dBFS · gain +9.5 dB · mls-00057
(emotional numbness, pain, sadness ·narration, monologue)je soulevai mon bonnet de fourrure aux trois quarts pelé et je commençai le signe de la croix en hanô an tad hag ar mab har ar spéred santel ailleurs la scène eût peut-être passé pour drôle et j'aurais probablement fait l'effet d'un singulier curé
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, pain, sadness; style: narration, monologue; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.7/10; 16.6s, FRENCH.
12709_14178_000706 · in -28.6 dBFS · gain +8.7 dB · mls-00057
(contemplation, longing, sadness · narration, monologue)mais là sur cette goélette solitaire dans l'infini silence et le vide infini il n'eût pas été du métier celui qui aurait eu le cœur de rire pour nous en vérité nous n'y pensions guère j'étais très grave et s'il faut l'avouer un peu ému comme du reste chaque fois qu'il m'est arrivé d'officier de la sorte
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, longing, sadness; style: narration, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.1/10; 18.7s, FRENCH.
12709_14178_000026 · in -28.2 dBFS · gain +8.2 dB · mls-00056
This chain comes from the two-sided rule: it only counts if both emotions move — Shame down and Emotional Numbness up — by at least 0.25 each.
The chain starts with Emotional Numbness clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.36.
At the same time Shame goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.09, then +0.20, then -0.01, then +0.08 — not a clean run: step 3 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 73 s · french · mls
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.850 before conversion and 0.870 after — it rose by 0.020. Neighbour-to-neighbour the worst pair went 0.791 → 0.870. (The earlier render, with segment 1 left raw, scores 0.783 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.363 in the original and +0.443 after conversion — 122 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Shame, -0.293 became -0.365.
Quality. Mean predicted overall quality across the segments went 3.18 → 3.25 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.850 → 0.870+0.020identity cos neighbours 0.791 → 0.870d_b rescored +0.363 → +0.443d_a rescored -0.293 → -0.365d_a mined -0.293d_b mined 0.363min_cos_consec (site) 0.8534min_cos_anchor (site) 0.8574dataset mlslang frenchspeaker 3698total 71.6schain gain +1.8 dBseam step 1.1 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an elderly somewhat feminine voice · balanced body, slightly relaxed, clear, light breath
(shame, triumph, malevolence malice · measured, very low-energy, fairly steady, ASMR)comme je ne voulais point parler haut je fis quatre ou cinq pas en retour de ceux que j'avais faits en avant j'allongeai les mains j'appelai avec précaution adieu la compagnie ils m'avaient laissé tout seul
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; clear, little disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as shame, triumph, malevolence malice; style: ASMR, whispered; average recording, no background noise; genuineness 1.1/6; vocal-burst blend 3.9/10; 12.7s, FRENCH.
3698_1902_000596 · in -28.1 dBFS · gain +8.1 dB · mls-00063
(fear, malevolence malice, anger· measured, normally alert, fairly steady, monologue)je pensai que n'étant pas bien loin de l'entrée je les rattraperais dedans ou dehors je marchai donc plus vite et avec plus d'assurance et repassai l'arcade par où j'étais entré pour regarder et chercher tout le long de la rouette aux anglais mais il était arrivé de mes camarades comme des sonneurs il semblait que la terre les eût dévorés
full caption & clip details
A young adult somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as fear, malevolence malice, anger; style: monologue, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 3.5/10; 20.0s, FRENCH.
3698_1902_000235 · in -28.0 dBFS · gain +8.0 dB · mls-00063
(shame, fear, malevolence malice ·normal-paced, normally alert, fairly steady, monologue)j'eus comme un moment de malefièvre en songeant qu'il me fallait tout abandonner ou rentrer dans ces maudites cavernes et m'y trouver tout seul aux prises avec les embûches et les frayeurs qui y attendaient joseph
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as shame, fear, malevolence malice; style: monologue, narration; good recording, no background noise; mildly explicit content; genuineness 0.2/6; vocal-burst blend 0.7/10; 11.5s, FRENCH.
3698_1902_000244 · in -27.0 dBFS · gain +7.0 dB · mls-00063
(fear, contemplation, longing·measured, subdued, steady, monologue)mais je me demandai si dans le cas où il ne s'agirait que de lui je me retirerais tranquillement de son danger mon âme de chrétien m'ayant répondu que non je demandai à mon coeur si l'amour de thérence n'était pas aussi solide en lui que l'amour du prochain dans ma conscience
full caption & clip details
A middle-aged somewhat feminine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as fear, contemplation, longing; style: monologue, whispered; average recording, no background noise; genuineness 0.4/6; vocal-burst blend 2.3/10; 16.0s, FRENCH.
3698_1902_000839 · in -26.8 dBFS · gain +6.8 dB · mls-00063
(emotional numbness, awe, contemplation · measured, very low-energy, fairly steady, monologue)et la réponse que j'en reçus me fit repasser l'arcade noire et vaseuse bien résolûment et courir dans le souterrain non pas aussi gai mais aussi prompt que si c'eût été à ma propre noce
full caption & clip details
A young adult somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness, awe, contemplation; style: monologue, storytelling; good recording, quiet background; genuineness 0.6/6; vocal-burst blend 2.6/10; 12.2s, FRENCH.
3698_1902_000130 · in -30.2 dBFS · gain +10.2 dB · mls-00063
This chain comes from the two-sided rule: it only counts if both emotions move — Contempt down and Relief up — by at least 0.25 each.
The chain starts with Relief clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.30.
At the same time Contempt goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.03, then +0.13, then +0.10, then +0.04 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 84 s · german · mls
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.940 before conversion and 0.937 after — it fell by 0.003. Neighbour-to-neighbour the worst pair went 0.915 → 0.930. (The earlier render, with segment 1 left raw, scores 0.843 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.304 in the original and +0.781 after conversion — 257 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contempt, -0.267 became -0.358.
Quality. Mean predicted overall quality across the segments went 3.01 → 3.27 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.940 → 0.937-0.003identity cos neighbours 0.915 → 0.930d_b rescored +0.304 → +0.781d_a rescored -0.267 → -0.358d_a mined -0.267d_b mined 0.304min_cos_consec (site) 0.9181min_cos_anchor (site) 0.9403dataset mlslang germanspeaker 3731total 82.6schain gain +1.9 dBseam step 3.1 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(contempt, disgust, malevolence malice · almost no disfluency, moderate pitch range, formal, monologue)unter diesen verschiedenen gesträuchen die so groß sind wie die bäume der gemäßigten zone und unter ihrem feuchten schatten befanden sich massenweis wahre gebüsche lebendiger pflanzen hecken von pflanzenthieren und was die täuschung vollends beförderte
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, disgust, malevolence malice; style: formal, monologue; average recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.0/10; 14.3s, GERMAN.
3731_1614_000036 · in -27.7 dBFS · gain +7.7 dB · mls-00017
(jealousy and envy, anger, teasing· almost no disfluency, moderate pitch range, formal, narration)die mückenfische flogen von zweig zu zweig gleich einem schwarm colibris während andere gleich einem trupp becassinen unter unseren schritten aufzufliegen schienen gegen ein uhr gab der kapitän nemo das zeichen zum halt ich meines theils war wohl zufrieden damit und wir streckten uns nieder
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, anger, teasing; style: formal, narration; average recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.0/10; 17.8s, GERMAN.
3731_1614_000019 · in -29.2 dBFS · gain +9.2 dB · mls-00017
(contempt, malevolence malice, emotional numbness· almost no disfluency, moderate pitch range, formal, narration)dieses ausruhen schien mir köstlich nur mangelte uns die unterhaltung denn das anreden war so unmöglich als das erwidern ich näherte nur meinen dicken kupferkopf dem conseil's ich sah bei diesem wackern jungen die augen vor befriedigung glänzen und um es kund zu geben machte er in seiner schale höchst komische bewegungen
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, malevolence malice, emotional numbness; style: formal, narration; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.0/10; 18.8s, GERMAN.
3731_1614_000035 · in -26.3 dBFS · gain +6.3 dB · mls-00017
(fear, infatuation, awe· almost no disfluency, moderate pitch range, narration, formal)nach einem vierstündigen spaziergang war ich sehr erstaunt daß ich nicht heftigen hunger empfand woher diese stimmung des magens kam konnte ich nicht sagen dagegen spürte ich eine unüberwindliche neigung zum schlafen wie das bei allen tauchern der fall ist daher schlossen sich auch alsbald meine augen hinter ihrem dichten glas
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, infatuation, awe; style: narration, formal; good recording, quiet background; genuineness 0.5/6; vocal-burst blend 0.0/10; 18.4s, GERMAN.
3731_1614_000084 · in -27.4 dBFS · gain +7.4 dB · mls-00017
(relief, fear, triumph·no disfluency, fairly narrow pitch, formal, whispered)und ich sank in eine unwiderstehliche schlaftrunkenheit welche bisher nur durch die bewegung des gehens zu bekämpfen möglich war der kapitän nemo nebst seinem kräftigen genossen gaben uns hingestreckt im klaren wasser das beispiel zum schlafen
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, fear, triumph; style: formal, whispered; average recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.0/10; 14.2s, GERMAN.
3731_1614_000049 · in -28.2 dBFS · gain +8.2 dB · mls-00017
This chain comes from the two-sided rule: it only counts if both emotions move — Longing down and Relief up — by at least 0.25 each.
The chain starts with Relief clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.36.
At the same time Longing goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.08, then +0.02, then +0.04, then +0.22 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 75 s · dutch · mls
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.948 before conversion and 0.930 after — it fell by 0.018. Neighbour-to-neighbour the worst pair went 0.946 → 0.938. (The earlier render, with segment 1 left raw, scores 0.727 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.359 in the original and +0.293 after conversion — 82 % of the delta retained, which is most of it. On the other named axis, Longing, -0.329 became +0.031.
Quality. Mean predicted overall quality across the segments went 3.21 → 3.43 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.948 → 0.930-0.018identity cos neighbours 0.946 → 0.938d_b rescored +0.359 → +0.293d_a rescored -0.329 → +0.031d_a mined -0.329d_b mined 0.359min_cos_consec (site) 0.9642min_cos_anchor (site) 0.9624dataset mlslang dutchspeaker 1724total 73.1schain gain +1.1 dBseam step 0.5 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an elderly somewhat feminine voice · neutral-toned, fairly smooth, average recording, quiet background, slightly relaxed, clear, light breath
(longing, contemplation, sourness · measured, very low-energy, fairly steady, whispered)ook wil ik u wel bekennen ging zij voort op vertrouwelijken toon dat ik er wel eens over gedacht heb om ze bij een lombard te beleenen maar neen neen dat moet niet zijn de gravin van solms moet niet in woekeraars handen vallen
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as longing, contemplation, sourness; style: whispered, monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.9/10; 14.2s, DUTCH.
1724_2757_000973 · in -23.4 dBFS · gain +3.4 dB · mls-00104
(contempt, longing, malevolence malice· measured, very low-energy, fairly steady, narration)hij zweeg en scheen zich even te bedenken zoo uwe genade nog niet besluiten kan zich van enkelen te ontdoen maar ze alleen tijdelijk wil afstaan weet ik een juwelier een eerlijk en discreet man die zulke zaken op hunne rechte waarde zal schatten
full caption & clip details
An elderly feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contempt, longing, malevolence malice; style: narration, whispered; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.4/10; 15.1s, DUTCH.
1724_2757_002127 · in -24.8 dBFS · gain +4.8 dB · mls-00104
(longing, contemplation, fear· measured, normally alert, steady, monologue)vertrouw aan mij wat gij meent te konnen missen ben ik eerst onderricht hoeveel dat te zamen uitmaakt dan is mijn broeder de schepen dirk jansz graswinckel de man die u daarop de verlangde som zal voorschieten zonder interest
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing, contemplation, fear; style: monologue, narration; average recording, quiet background; genuineness 0.8/6; vocal-burst blend 0.0/10; 12.5s, DUTCH.
1724_2757_000108 · in -24.7 dBFS · gain +4.7 dB · mls-00104
(anger, contemplation ·normal-paced, normally alert, fairly steady, narration)hij is met den handel rijk geworden heeft geene kinderen en is steeds geneigd op mijne voorspraak iemand in stilte te verplichten maar hij verlangt zeker onderpand voor de teruggave zij die op langen tijd uitgesteld en in dat zwak moeten wij hem met uw goedvinden te gemoet komen kan uwe genade in dit voorstel treden
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger, contemplation; style: narration, whispered; average recording, quiet background; genuineness 0.9/6; vocal-burst blend 0.0/10; 17.9s, DUTCH.
1724_2757_002562 · in -24.7 dBFS · gain +4.7 dB · mls-00104
(relief, sourness, thankfulness gratitude·measured, very low-energy, fairly steady, whispered)ik zou u als mijn reddenden engel beschouwen zoo gij dit voor mij kondet verkrijgen ik wenschte u de helpende hand toe te reiken voor beter dan dit wellieve vrouwe doch in de moeielijkheid van t oogenblik moet nu allereerst worden voorzien
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; clear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as relief, sourness, thankfulness gratitude; style: whispered, narration; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 0.9/10; 14.2s, DUTCH.
1724_2757_000262 · in -24.7 dBFS · gain +4.7 dB · mls-00104
This chain comes from the two-sided rule: it only counts if both emotions move — Sourness down and Concentration up — by at least 0.25 each.
The chain starts with Concentration around average — 0.57, higher than 57 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.36.
At the same time Sourness goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.12, then +0.24 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 41 s · italian · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.912 before conversion and 0.910 after — it fell by 0.002. Neighbour-to-neighbour the worst pair went 0.925 → 0.896. (The earlier render, with segment 1 left raw, scores 0.778 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.365 in the original and +0.510 after conversion — 140 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Sourness, -0.251 became -0.477.
Quality. Mean predicted overall quality across the segments went 2.90 → 3.33 (+0.44) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.912 → 0.910-0.002identity cos neighbours 0.925 → 0.896d_b rescored +0.365 → +0.510d_a rescored -0.251 → -0.477d_a mined -0.255d_b mined 0.364min_cos_consec (site) 0.9273min_cos_anchor (site) 0.9173dataset mlslang italianspeaker 10446total 40.5schain gain +1.9 dBseam step 0.9 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, balanced body, average recording, quiet background, normally alert, slightly relaxed, fairly steady, light breath
(sourness, contentment, jealousy and envy · measured, some disfluency, average clarity, didactic)meno grave all'aspetto più elegante nel portamento ma pur sempre severo e rispondente a quell'immagine di dignità e di forza che non dovrebbe scompagnarsi mai dall'idea dell'uomo
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, contentment, jealousy and envy; style: didactic, monologue; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 3.4/10; 11.7s, ITALIAN.
10446_10415_000708 · in -25.8 dBFS · gain +5.8 dB · mls-00065
(contentment, affection, pride·normal-paced, some disfluency, average clarity, monologue)egli frattanto non vedeva più il monachino ma una bella e graziosa fanciulla l'aria birichina dello scolaro in vacanze non c'era più ma l'aspetto della donna che sente e che pensa rendeva il suo volto anche più attraente che non fosse da prima
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contentment, affection, pride; style: monologue, authoritative; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 1.0/10; 14.5s, ITALIAN.
10446_10415_001555 · in -26.0 dBFS · gain +6.0 dB · mls-00066
(concentration·measured, frequent disfluency, somewhat unclear, monologue)come balbettò ella avvicinandosi signorina eccomi qua rispose egli dissimulando con un profondo inchino la sua profonda commozione seguì una scena muta di forse un minuto il solito minuto che parve un secolo
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, didactic; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 1.6/10; 14.7s, ITALIAN.
10446_10415_001216 · in -25.9 dBFS · gain +5.9 dB · mls-00066
This chain comes from the two-sided rule: it only counts if both emotions move — Thankfulness Gratitude down and Shame up — by at least 0.25 each.
The chain starts with Shame clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.31.
At the same time Thankfulness Gratitude goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.03, then +0.09 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 65 s · polish · mls
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.869 before conversion and 0.893 after — it rose by 0.024. Neighbour-to-neighbour the worst pair went 0.903 → 0.915. (The earlier render, with segment 1 left raw, scores 0.865 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.312 in the original and +0.180 after conversion — 58 % of the delta retained. On the other named axis, Thankfulness Gratitude, -0.261 became -0.278.
Quality. Mean predicted overall quality across the segments went 3.15 → 3.45 (+0.31) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.869 → 0.893+0.024identity cos neighbours 0.903 → 0.915d_b rescored +0.312 → +0.180d_a rescored -0.261 → -0.278d_a mined -0.261d_b mined 0.312min_cos_consec (site) 0.9378min_cos_anchor (site) 0.9144dataset mlslang polishspeaker 6892total 64.1schain gain +0.8 dBseam step 1.0 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a child masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(thankfulness gratitude, infatuation, contentment · no disfluency, average clarity, narration, monologue)tu opowiada ący uśmiechnął się z gorzką ironią gdybym miał żonę zbytnicę i trzech synów utrac uszów nie wiem czy by mię tyle kosztowali co ten eden człowiek
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, infatuation, contentment; style: narration, monologue; good recording, quiet background; genuineness 0.3/6; vocal-burst blend 0.0/10; 10.3s, POLISH.
6892_10462_000754 · in -23.0 dBFS · gain +3.0 dB · mls-00106
(contentment, infatuation, affection·some disfluency, average clarity, narration, monologue)na cóż on tak wydawał zawołałyśmy obie narkotyki pieniądz bóg raczy wiedzieć na co zawsze mówił że mu potrzeba ogromnych sum na doświadczenia naukowe i rzeczywiście prowadził e z wielkim nakładem ale co eszcze nierównie więce pochłaniało to ego codzienne niewyczerpane fantaz
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contentment, infatuation, affection; style: narration, monologue; good recording, quiet background; genuineness 0.7/6; vocal-burst blend 3.8/10; 17.1s, POLISH.
6892_10462_000999 · in -23.2 dBFS · gain +3.2 dB · mls-00106
(contemplation, sexual lust, shame· some disfluency, average clarity, narration, storytelling)i tak na przykład na droższe tytunie nazywał mdłą trawą emu potrzebny był haszysz ów narkotyk indy ski którego tak trudno u nas dostać a ednak musiałem go ciągle sprowadzać za ba eczne sumy bo kilka dni przebytych bez haszyszu wtrącało go w ponurość i
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, sexual lust, shame; style: narration, storytelling; average recording, quiet background; genuineness 0.4/6; vocal-burst blend 1.9/10; 18.1s, POLISH.
6892_10462_001019 · in -23.9 dBFS · gain +3.9 dB · mls-00106
(shame, pain, contemplation · some disfluency, clear, narration, storytelling)takich niebezpiecznych nawyknień miał mnóstwo w każdym kra u uczył się nowych a zwiedziliśmy kra ów niemało namówił mię na podróże na które chętnie przystałem bo mo e rodzinne strony obrzydły mi po kilku latach dla nieustannych dokuczań akich doznawałem od obrażonych kuzynów i całe rodziny
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, pain, contemplation; style: narration, storytelling; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 3.4/10; 19.2s, POLISH.
6892_10462_000679 · in -23.0 dBFS · gain +3.0 dB · mls-00106
This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Bitterness up — by at least 0.25 each.
The chain starts with Bitterness clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.30.
At the same time Emotional Numbness goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.20, then -0.06, then +0.16 — not a clean run: step 2 moves back the other way by 0.06 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 56 s · german · mls
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.926 before conversion and 0.914 after — it fell by 0.013. Neighbour-to-neighbour the worst pair went 0.950 → 0.926. (The earlier render, with segment 1 left raw, scores 0.834 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Bitterness moved +0.295 in the original and +0.141 after conversion — 48 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Emotional Numbness, -0.286 became -0.241.
Quality. Mean predicted overall quality across the segments went 3.25 → 3.33 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.926 → 0.914-0.013identity cos neighbours 0.950 → 0.926d_b rescored +0.295 → +0.141d_a rescored -0.286 → -0.241d_a mined -0.286d_b mined 0.295min_cos_consec (site) 0.9522min_cos_anchor (site) 0.9470dataset mlslang germanspeaker 11927total 54.9schain gain +3.6 dBseam step 1.4 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, normally alert, slightly relaxed, fairly steady
(emotional numbness, sadness, relief · measured, almost no disfluency, monologue, narration)denn obzwar er viel von dem prediger beschworen wurde auch männiglich in der kirche auf die kniee gefallen und fleißig und andächtig gebetet so hat er doch mit dem austreiben nichts als spott und kurzweil getrieben
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, sadness, relief; style: monologue, narration; average recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 14.6s, GERMAN.
11927_12408_001302 · in -29.2 dBFS · gain +9.2 dB · mls-00011
(relief, sadness, jealousy and envy· measured, some disfluency, monologue, narration)so hat er oft gesagt ja er wolle weichen er müsse auch wohl räumen aber er hat allerlei gefordert ihm zu erlauben daß er es mitnehmen dürfe wann ihm dann nun das eine abgeschlagen wurde so hatte er gleich das andere bei der hand
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, sadness, jealousy and envy; style: monologue, narration; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.0/10; 16.4s, GERMAN.
11927_12408_000412 · in -26.7 dBFS · gain +6.7 dB · mls-00010
(pride, relief, malevolence malice·slow, frequent disfluency, monologue, narration)es stand einer in der kirchen der den hut aufbehalten hatte da forderte er von dem prediger daß er den hut dem menschen vom kopfe nehmen und mit sich führen dürfe
full caption & clip details
A child masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, relief, malevolence malice; style: monologue, narration; average recording, quiet background; genuineness 0.6/6; vocal-burst blend 0.0/10; 11.8s, GERMAN.
11927_12408_001071 · in -27.1 dBFS · gain +7.1 dB · mls-00010
(bitterness, distress, longing·measured, almost no disfluency, didactic, monologue)aber der prediger trug mit recht sorge wenn er ihm den hut gestattet so hätten mit dem hute auch haut und haar davongehen müssen letztlich aber als er vermerkte daß seine zeit verflossen
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as bitterness, distress, longing; style: didactic, monologue; average recording, quiet background; genuineness 0.7/6; vocal-burst blend 0.0/10; 12.7s, GERMAN.
11927_12408_001540 · in -27.3 dBFS · gain +7.3 dB · mls-00011
Fear ↓ / Jealousy and Envy ↑identity −0.02emotion 60 % c-mls-AB2 · #13
This chain comes from the two-sided rule: it only counts if both emotions move — Fear down and Jealousy and Envy up — by at least 0.25 each.
The chain starts with Jealousy and Envy clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.28.
At the same time Fear goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.07 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 45 s · german · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.896 before conversion and 0.873 after — it fell by 0.023. Neighbour-to-neighbour the worst pair went 0.902 → 0.893. (The earlier render, with segment 1 left raw, scores 0.753 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Jealousy and Envy moved +0.279 in the original and +0.167 after conversion — 60 % of the delta retained. On the other named axis, Fear, -0.287 became -0.787.
Quality. Mean predicted overall quality across the segments went 3.01 → 3.35 (+0.33) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.896 → 0.873-0.023identity cos neighbours 0.902 → 0.893d_b rescored +0.279 → +0.167d_a rescored -0.287 → -0.787d_a mined -0.287d_b mined 0.279min_cos_consec (site) 0.9281min_cos_anchor (site) 0.9121dataset mlslang germanspeaker 252total 44.4schain gain +3.3 dBseam step 1.3 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, balanced body, no background noise, subdued, slightly relaxed, almost no disfluency, clear
(malevolence malice, fear, sadness · slow, steady, fairly narrow pitch, narration)und der alte schwarzkünstler schien dem nichts troz bieten zu wollen so sah er aus als er den teufel bannte sagte die wahrsagerin
full caption & clip details
A middle-aged masculine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, fear, sadness; style: narration, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.6/10; 11.5s, GERMAN.
252_1552_000553 · in -24.4 dBFS · gain +4.4 dB · mls-00016
(malevolence malice, bitterness, contempt· slow, fairly steady, fairly narrow pitch, storytelling)nur haben sie ihm nachher die hände gefaltet daß er hier unten wider willen beten muß und warum betet er denn fragte ich zornig da drüben über uns im himmelssee funkeln und schwimmen zwar unzählige sterne
full caption & clip details
A middle-aged masculine voice; delivery is subdued, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, rough, balanced body; clear, almost no disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, bitterness, contempt; style: storytelling, narration; average recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.1/10; 15.4s, GERMAN.
252_1552_000588 · in -21.0 dBFS · gain +1.0 dB · mls-00016
(jealousy and envy, pride, contempt ·measured, fairly steady, moderate pitch range, storytelling)aber wenn es welten sind wie viele kluge köpfe behaupten so giebt es auch schädel auf ihnen und würmer wie hier unten das geht so fort durch die ganze unermeßlichkeit und der baseler todtentanz wird dadurch nur um so lustiger und wilder und der ballsaal größer
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, pride, contempt; style: storytelling, monologue; average recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.0/10; 18.0s, GERMAN.
252_1552_000213 · in -21.2 dBFS · gain +1.2 dB · mls-00016
This chain comes from the two-sided rule: it only counts if both emotions move — Shame down and Emotional Numbness up — by at least 0.25 each.
The chain starts with Emotional Numbness clearly present — 0.59, higher than 59 % of clips in this corpus — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.29.
At the same time Shame goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.09, then +0.19 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 40 s · polish · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.940 before conversion and 0.951 after — it rose by 0.011. Neighbour-to-neighbour the worst pair went 0.946 → 0.953. (The earlier render, with segment 1 left raw, scores 0.890 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.287 in the original and +0.233 after conversion — 81 % of the delta retained, which is most of it. On the other named axis, Shame, -0.254 became -0.172.
Quality. Mean predicted overall quality across the segments went 3.18 → 3.44 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.940 → 0.951+0.011identity cos neighbours 0.946 → 0.953d_b rescored +0.287 → +0.233d_a rescored -0.254 → -0.172d_a mined -0.271d_b mined 0.287min_cos_consec (site) 0.9484min_cos_anchor (site) 0.9402dataset mlslang polishspeaker 6892total 39.3schain gain +0.5 dBseam step 1.2 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(shame, infatuation, longing · average clarity, narration, storytelling)pokoju widzenie znikło wokulski ocknął się był znowu tylko człowiekiem zbolałym i słabym ale w jego duszy huczał jakiś potężny głos niby echo kwietniowej burzy grzmotami zapowiadającej wiosnę i zmartwychwstanie
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, infatuation, longing; style: narration, storytelling; average recording, quiet background; genuineness 0.9/6; vocal-burst blend 1.8/10; 13.0s, POLISH.
6892_9500_000322 · in -25.1 dBFS · gain +5.1 dB · mls-00112
(shame, contentment, sadness·clear, narration, monologue)pierwszego czerwca odwiedził go szlangbaum wszedł zakłopotany ale przypatrzywszy się wokulskiemu nabrał otuchy nie odwiedzałem cię do tej pory zaczął bo wiem żeś był niezdrów i nie chciałeś nikogo widywać no ale dzięki bogu już wszystko przeszło
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, contentment, sadness; style: narration, monologue; average recording, quiet background; genuineness 0.3/6; vocal-burst blend 3.8/10; 15.1s, POLISH.
6892_9500_000116 · in -25.6 dBFS · gain +5.6 dB · mls-00112
(average clarity, monologue, narration)kręcił się na krześle i zpod oka rzucał spojrzenia na pokój może spodziewał się znaleść w nim więcej nieładu masz jaki interes zapytał go wokulski nie tyle interes ile propozycyą
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, quiet background; genuineness 0.8/6; vocal-burst blend 2.1/10; 11.6s, POLISH.
6892_9500_000375 · in -25.9 dBFS · gain +5.9 dB · mls-00112
This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Interest up — by at least 0.25 each.
The chain starts with Interest below average — 0.40, lower than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.59.
At the same time Emotional Numbness goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.47 (lower than 53 % of clips in this corpus), a change of -0.48. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.19, then +0.19, then +0.20 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 62 s · polish · mls
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.933 before conversion and 0.914 after — it fell by 0.019. Neighbour-to-neighbour the worst pair went 0.938 → 0.915. (The earlier render, with segment 1 left raw, scores 0.817 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.590 in the original and +0.553 after conversion — 94 % of the delta retained, which is essentially all of it. On the other named axis, Emotional Numbness, -0.481 became -0.728.
Quality. Mean predicted overall quality across the segments went 3.15 → 3.43 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.933 → 0.914-0.019identity cos neighbours 0.938 → 0.915d_b rescored +0.590 → +0.553d_a rescored -0.481 → -0.728d_a mined -0.481d_b mined 0.590min_cos_consec (site) 0.9531min_cos_anchor (site) 0.9421dataset mlslang polishspeaker 6892total 60.7schain gain +1.2 dBseam step 0.6 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a child masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert, slightly relaxed, some disfluency
(emotional numbness, disgust · normal-paced, steady, average clarity, narration)gdzie ona się znalazła tam obok niej wszystko bladło inne kobiety były jej tłem a mężczyźni niewolnikami i to wszystko przeszło i dziś w tym salonie chłodno ciemno i pusto
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, disgust; style: narration, monologue; good recording, quiet background; genuineness 0.1/6; vocal-burst blend 0.0/10; 12.3s, POLISH.
6892_8764_001950 · in -25.1 dBFS · gain +5.2 dB · mls-00111
(bitterness, triumph, pride·fast, fairly steady, slurred, monologue)jest tylko ona i niewidzialny pająk smutku który zawsze zasnuwa szarą siecią te miejsca gdzie byliśmy szczęśliwi i zkąd szczęście uciekło już uciekło panna izabela wyłamywała sobie palce ażeby pohamować się od łez których wstyd jej było nawet w pustce i w nocy
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as bitterness, triumph, pride; style: monologue, narration; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 5.2/10; 16.7s, POLISH.
6892_8764_000938 · in -25.5 dBFS · gain +5.5 dB · mls-00111
(pride, contemplation, sadness·normal-paced, fairly steady, average clarity, narration)wszyscy ją opuścili z wyjątkiem hrabiny karolowej która kiedy wezbrał jej zły humor przychodziła tu i szeroko zasiadłszy na kanapie prawiła wśród westchnień tak droga belciu musisz przyznać że popełniłaś kilka błędów nie
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, contemplation, sadness; style: narration, monologue; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 4.2/10; 14.1s, POLISH.
6892_8764_001708 · in -24.4 dBFS · gain +4.4 dB · mls-00111
(interest, pride, bitterness· normal-paced, fairly steady, average clarity, monologue)nie mówię o wiktorze emanuelu bo tamto był przelotny kaprys króla trochę liberalnego i zresztą bardzo zadłużonego na takie stosunki trzeba mieć więcej nie powiem taktu ale doświadczenia ciagnęłaciągnęła hrabina skromnie spuszczając powieki ale wypuścić czy jeżeli chcesz odrzucić
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, pride, bitterness; style: monologue, storytelling; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 6.9/10; 18.2s, POLISH.
6892_8764_000933 · in -24.6 dBFS · gain +4.6 dB · mls-00111
This chain comes from the two-sided rule: it only counts if both emotions move — Affection down and Triumph up — by at least 0.25 each.
The chain starts with Triumph around average — 0.51, higher than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.47.
At the same time Affection goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.23 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.76 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.76 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.76, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 37 s · italian · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.737 before conversion and 0.684 after — it fell by 0.053. Neighbour-to-neighbour the worst pair went 0.769 → 0.816. (The earlier render, with segment 1 left raw, scores 0.742 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Triumph moved +0.475 in the original and +0.550 after conversion — 116 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Affection, -0.398 became -0.387.
Quality. Mean predicted overall quality across the segments went 3.15 → 3.24 (+0.09) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.737 → 0.684-0.053identity cos neighbours 0.769 → 0.816d_b rescored +0.475 → +0.550d_a rescored -0.398 → -0.387d_a mined -0.398d_b mined 0.475min_cos_consec (site) 0.7584min_cos_anchor (site) 0.7578dataset mlslang italianspeaker 4974total 35.8schain gain +1.0 dBseam step 0.9 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an elderly somewhat masculine voice · neutral-toned, average recording
(affection, confusion, awe · slow, very low-energy, tense, dramatic)lei tu carlino lei tu ma come io oh dio ma che fui crudele lo riconosco
full caption & clip details
An elderly somewhat masculine voice; delivery is very low-energy, slow, tense, variable; timbre is neutral-toned, bright, gravelly, slightly thin; slurred, some disfluency, very wide pitch range, audible breath; affect is elated, very submissive, very vulnerable; reads as affection, confusion, awe; style: dramatic, casual; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 4.6/10; 10.0s, ITALIAN.
4974_5568_000394 · in -27.8 dBFS · gain +7.8 dB · mls-00075
(shame, embarrassment, distress·measured, normally alert, slightly relaxed, monologue)e tanto più mi dolgo della mia crudeltà in quanto quel poverino dovette credere senza dubbio ch'io avessi voluto prendermi il gusto di smascherarlo di fronte al paese con quella burla
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, embarrassment, distress; style: monologue, narration; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 3.3/10; 12.5s, ITALIAN.
4974_5568_000144 · in -31.4 dBFS · gain +11.4 dB · mls-00075
(triumph, shame, relief· measured, very low-energy, slightly relaxed, narration)mentre ero più che sicuro della sua buona fede più che sicuro ormai d'essere stato uno sciocco a maravigliarmi tanto poiché io stesso avevo già sperimentato
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, thin; slurred, frequent disfluency, narrow pitch range, audible breath; affect is mildly negative, neutral stance, slightly guarded; reads as triumph, shame, relief; style: narration, whispered; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 2.5/10; 13.7s, ITALIAN.
4974_5568_000049 · in -32.5 dBFS · gain +12.6 dB · mls-00075
This chain comes from the two-sided rule: it only counts if both emotions move — Jealousy and Envy down and Emotional Numbness up — by at least 0.25 each.
The chain starts with Emotional Numbness clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.28.
At the same time Jealousy and Envy goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.00, then +0.12, then +0.00, then +0.16 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 87 s · dutch · mls
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.876 before conversion and 0.852 after — it fell by 0.024. Neighbour-to-neighbour the worst pair went 0.876 → 0.869. (The earlier render, with segment 1 left raw, scores 0.748 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.276 in the original and +0.261 after conversion — 94 % of the delta retained, which is essentially all of it. On the other named axis, Jealousy and Envy, -0.251 became -0.330.
Quality. Mean predicted overall quality across the segments went 2.98 → 3.47 (+0.49) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.876 → 0.852-0.024identity cos neighbours 0.876 → 0.869d_b rescored +0.276 → +0.261d_a rescored -0.251 → -0.330d_a mined -0.251d_b mined 0.276min_cos_consec (site) 0.8763min_cos_anchor (site) 0.8763dataset mlslang dutchspeaker 2450total 85.5schain gain +1.7 dBseam step 1.0 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · neutral-toned, slightly dark, average recording, quiet background, measured, normally alert, slightly relaxed, fairly steady
(jealousy and envy, disgust, concentration · some disfluency, slurred, audible breath, narration)de heer boreel is oud en ziekelijk hervatte van espenblad zonder die loftuiting op zijn uitzicht te beantwoorden en de heer van beuningen is niet rekkelijk en volgzaam genoeg naar den zin van uwe regeering gronden te over waarop de heer de
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, some disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, disgust, concentration; style: narration, cartoonish; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 2.3/10; 16.5s, DUTCH.
2450_10026_002346 · in -28.1 dBFS · gain +8.1 dB · mls-00084
(jealousy and envy, disgust, concentration · some disfluency, slurred, audible breath, narration)de heer boreel is oud en ziekelijk hervatte van espenblad zonder die loftuiting op zijn uitzicht te beantwoorden en de heer van beuningen is niet rekkelijk en volgzaam genoeg naar den zin van uwe regeering gronden te over waarop de heer de
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, some disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, disgust, concentration; style: narration, cartoonish; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 2.3/10; 16.5s, DUTCH.
2450_10026_002346 · in -28.1 dBFS · gain +8.1 dB · mls-00084
(anger, contempt, concentration ·frequent disfluency, somewhat unclear, audible breath, cartoonish)onder s hands aan de witt zoû kunnen te kennen geven dat een ander ambassadeur niet onwelkom zoû wezen aan het fransche hof wees overtuigd zeide d'estrades dat de heer de lionne zoowel als de koning mijn meester volkomen bekend zijn met de schatting waarin ik u houde
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, slightly dominant, fairly guarded; reads as anger, contempt, concentration; style: cartoonish, monologue; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.9/10; 19.5s, DUTCH.
2450_10026_000657 · in -29.0 dBFS · gain +9.0 dB · mls-00084
(anger, contempt, concentration · frequent disfluency, somewhat unclear, audible breath, cartoonish)onder s hands aan de witt zoû kunnen te kennen geven dat een ander ambassadeur niet onwelkom zoû wezen aan het fransche hof wees overtuigd zeide d'estrades dat de heer de lionne zoowel als de koning mijn meester volkomen bekend zijn met de schatting waarin ik u houde
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, slightly dominant, fairly guarded; reads as anger, contempt, concentration; style: cartoonish, monologue; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.9/10; 19.5s, DUTCH.
2450_10026_000657 · in -29.0 dBFS · gain +9.0 dB · mls-00084
(concentration, emotional numbness, malevolence malice·some disfluency, slurred, light breath, narration)in weêrwil van al zijn slimheid en menschenkennis was van espenblad toch genoeg door ydelheid verblind om deze woorden van den graaf voor een kompliment op te nemen en hy boog zich op de wellevendste wijze
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, emotional numbness, malevolence malice; style: narration, monologue; average recording, quiet background; genuineness 0.2/6; vocal-burst blend 0.9/10; 14.3s, DUTCH.
2450_10026_003309 · in -28.8 dBFS · gain +8.8 dB · mls-00084
This chain comes from the two-sided rule: it only counts if both emotions move — Anger down and Longing up — by at least 0.25 each.
The chain starts with Longing clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.30.
At the same time Anger goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.15 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 50 s · german · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.860 before conversion and 0.841 after — it fell by 0.020. Neighbour-to-neighbour the worst pair went 0.874 → 0.884. (The earlier render, with segment 1 left raw, scores 0.707 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.301 in the original and +0.982 after conversion — 327 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Anger, -0.271 became -0.092.
Quality. Mean predicted overall quality across the segments went 3.10 → 3.28 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.860 → 0.841-0.020identity cos neighbours 0.874 → 0.884d_b rescored +0.301 → +0.982d_a rescored -0.271 → -0.092d_a mined -0.271d_b mined 0.299min_cos_consec (site) 0.9364min_cos_anchor (site) 0.9145dataset mlslang germanspeaker 10148total 49.1schain gain +2.9 dBseam step 0.5 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, slightly relaxed, clear
(anger, fear, impatience and irritability · normal-paced, normally alert, fairly steady, monologue)ohne weitere umschweife erklärte er sich unverzüglich bereit für alle die bauern die auf eine so unglückliche weise gestorben waren aus eigener tasche die abgaben zu entrichten dieser vorschlag versetzte pljuschkin in höchstes erstaunen er glotzte ihn lange an und fragte zuletzt
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger, fear, impatience and irritability; style: monologue, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 19.5s, GERMAN.
10148_11442_003518 · in -28.8 dBFS · gain +8.8 dB · mls-00008
(fear, confusion, sadness· normal-paced, normally alert, fairly steady, didactic)waren sie vielleicht im militärdienst väterchen nein entgegnete tschitschikow nicht ohne list ich war nur im zivildienste im zivildienste wiederholte pljuschkin und begann die lippen zu bewegen als ob er etwas kaute
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as fear, confusion, sadness; style: didactic, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.1/10; 18.6s, GERMAN.
10148_11442_003581 · in -28.8 dBFS · gain +8.8 dB · mls-00008
(longing, sadness, malevolence malice·measured, subdued, steady, didactic)wie ist es nun das wäre doch ein schaden für sie um ihnen ein vergnügen zu bereiten bin ich auch bereit den schaden auf mich zu nehmen
full caption & clip details
A middle-aged somewhat feminine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing, sadness, malevolence malice; style: didactic, whispered; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.7/10; 11.4s, GERMAN.
10148_11442_003849 · in -29.6 dBFS · gain +9.6 dB · mls-00008
This chain comes from the two-sided rule: it only counts if both emotions move — Sourness down and Infatuation up — by at least 0.25 each.
The chain starts with Infatuation clearly present — 0.73, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.26.
At the same time Sourness goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.74 (higher than 74 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.18, then +0.05, then +0.03 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.77 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.77 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.77, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 55 s · spanish · mls
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.786 before conversion and 0.823 after — it rose by 0.038. Neighbour-to-neighbour the worst pair went 0.786 → 0.834. (The earlier render, with segment 1 left raw, scores 0.722 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.262 in the original and +0.018 after conversion — 7 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Sourness, -0.260 became -0.597.
Quality. Mean predicted overall quality across the segments went 2.75 → 3.26 (+0.52) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.786 → 0.823+0.038identity cos neighbours 0.786 → 0.834d_b rescored +0.262 → +0.018d_a rescored -0.260 → -0.597d_a mined -0.260d_b mined 0.262min_cos_consec (site) 0.7742min_cos_anchor (site) 0.7742dataset mlslang spanishspeaker 8882total 54.0schain gain +1.4 dBseam step 3.4 dBcrossfades 150/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · balanced body, measured
(sourness, malevolence malice, contentment · normally alert, neutral tension, moderately variable, storytelling)gritó con insólito buen humor es verdaderamente necesario que les dé usted un descanso á sus caballos y que entre á tomar un vaso de vino y á felicitarme
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is warm, slightly dark, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as sourness, malevolence malice, contentment; style: storytelling, cartoonish; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 5.3/10; 13.1s, SPANISH.
8882_10372_000146 · in -21.7 dBFS · gain +1.7 dB · mls-00025
(sourness, emotional numbness, bitterness· normally alert, slightly relaxed, fairly steady, monologue)después de lo que había oído decir sobre la manera como erankland trataba á su hija mis sentimientos hacia él estaban lejos de ser amistosos pero deseaba despachar á perkins y al
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, emotional numbness, bitterness; style: monologue, didactic; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 4.2/10; 14.2s, SPANISH.
8882_10372_000017 · in -22.2 dBFS · gain +2.2 dB · mls-00025
(contentment, bitterness, triumph·subdued, slightly relaxed, fairly steady, monologue)para queaarme solo y la ocasión era buena bajé del carruaje y envió un recado á sir enrique haciéndole baber que regresaría á pie á tiempo para la comida después seguí á frankland hasta el comedor de su casa
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contentment, bitterness, triumph; style: monologue, didactic; average recording, quiet background; genuineness 0.6/6; vocal-burst blend 3.5/10; 17.0s, SPANISH.
8882_10372_000250 · in -22.4 dBFS · gain +2.4 dB · mls-00025
(infatuation, thankfulness gratitude, elation·normally alert, slightly relaxed, moderately variable, storytelling)hoy es un gran día para mí señor uno de los verdaderos días de fiesta de mi vida exclamó entre risitas ahogadas
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; very clear, some disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as infatuation, thankfulness gratitude, elation; style: storytelling, narration; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 2.6/10; 10.2s, SPANISH.
8882_10372_000554 · in -21.7 dBFS · gain +1.7 dB · mls-00025
This chain comes from the two-sided rule: it only counts if both emotions move — Fatigue Exhaustion down and Affection up — by at least 0.25 each.
The chain starts with Affection clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.38.
At the same time Fatigue Exhaustion goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.67 (higher than 67 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 35 s · french · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.945 before conversion and 0.903 after — it fell by 0.042. Neighbour-to-neighbour the worst pair went 0.937 → 0.924. (The earlier render, with segment 1 left raw, scores 0.830 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Affection moved +0.382 in the original and +0.362 after conversion — 95 % of the delta retained, which is essentially all of it. On the other named axis, Fatigue Exhaustion, -0.324 became -0.719.
Quality. Mean predicted overall quality across the segments went 3.02 → 3.21 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.945 → 0.903-0.042identity cos neighbours 0.937 → 0.924d_b rescored +0.382 → +0.362d_a rescored -0.324 → -0.719d_a mined -0.324d_b mined 0.382min_cos_consec (site) 0.9414min_cos_anchor (site) 0.9592dataset mlslang frenchspeaker 123total 34.6schain gain +1.9 dBseam step 0.2 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, slightly relaxed, clear, light breath
(fatigue exhaustion, disgust, distress · fast, normally alert, fairly steady, monologue)comme il n'en pouvoit plus de fatigue il s'endormit aprés s'estre reposé quelque temps et vint à ronfler si effroyablement que les pauvres enfans n'eurent pas moins de peur que quand il tenoit son grand couteau pour leur couper la gorge
full caption & clip details
A child masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, disgust, distress; style: monologue, authoritative; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.0/10; 13.4s, FRENCH.
123_131_000192 · in -31.8 dBFS · gain +11.8 dB · mls-00052
(fear, emotional numbness, distress · fast, normally alert, fairly steady, monologue)le petit poucet en eut moins de peur et dit à ses freres de s'enfuir promptement à la maison pendant que l'ogre dormoit bien fort et qu'ils ne se missent point en peine de luy
full caption & clip details
A child masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, emotional numbness, distress; style: monologue, authoritative; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.0s, FRENCH.
123_131_000087 · in -33.8 dBFS · gain +13.8 dB · mls-00052
(affection, malevolence malice, contentment·measured, subdued, steady, monologue)ils crurent son conseil et gagnerent viste la maison le petit poucet s'estant approché de l'ogre lui tira doucement ses bottes et les mit aussi
full caption & clip details
A child masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as affection, malevolence malice, contentment; style: monologue, cartoonish; average recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.6s, FRENCH.
123_131_000193 · in -35.1 dBFS · gain +15.1 dB · mls-00052