c-mls-PXR — voice-corrected

Corpus mls in isolation, rule PXR.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_c-mls-PXR.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
57segments re-voiced
0.917 → 0.908median worst-to-anchor identity cosine
66 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Relief ↓  /  Angeridentity −0.01 emotion 81 %   c-mls-PXR · #1

This chain comes from the proxy rule: the same two-sided test as above, but because Anger is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Anger around average — 0.58, higher than 58 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.42.

At the same time Relief goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.63 (higher than 63 % of clips in this corpus), a change of -0.35. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.19, then +0.00, then +0.23, then +0.00 — a plateau around step 2, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 66 s · dutch · mls

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.947 before conversion and 0.938 after — it fell by 0.010. Neighbour-to-neighbour the worst pair went 0.938 → 0.910. (The earlier render, with segment 1 left raw, scores 0.771 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Anger moved +0.416 in the original and +0.339 after conversion — 81 % of the delta retained, which is most of it. On the other named axis, Relief, -0.350 became -0.218.

Quality. Mean predicted overall quality across the segments went 3.19 → 3.45 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.947 → 0.938 -0.010identity cos neighbours 0.938 → 0.910d_b rescored +0.416 → +0.339d_a rescored -0.350 → -0.218d_a mined -0.350d_b mined 0.416min_cos_consec (site) 0.9418min_cos_anchor (site) 0.9533dataset mlslang dutchspeaker 1724total 65.1schain gain +3.1 dBseam step 1.4 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a child somewhat feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, measured, slightly relaxed, light breath
(relief, fatigue exhaustion · subdued, fairly steady, frequent disfluency, monologue) of vervolgde bol de septer van ageen dan iets anders ware geweest dan een stok ofschoon ik beken nergens gelezen te hebben dat hij zich met een tang behielp voegde hij er halfluid bij
full caption & clip details
A child somewhat feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as relief, fatigue exhaustion; style: monologue, narration; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.0/10; 13.5s, DUTCH.
1724_10176_000469 · in -28.5 dBFS · gain +8.5 dB · mls-00085
(jealousy and envy, contempt, sourness · normally alert, fairly steady, some disfluency, monologue) maar hervatte hij zich herinnerende welken plicht hij als gastheer te vervullen had wij zullen dat punt te zijner tijd behandelen je vergeet geheel katuil dat je niet alleen komt
full caption & clip details
A young adult somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as jealousy and envy, contempt, sourness; style: monologue, storytelling; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 0.2/10; 11.9s, DUTCH.
1724_10176_000285 · in -26.5 dBFS · gain +6.5 dB · mls-00085
(jealousy and envy, contempt, sourness · normally alert, fairly steady, some disfluency, monologue) maar hervatte hij zich herinnerende welken plicht hij als gastheer te vervullen had wij zullen dat punt te zijner tijd behandelen je vergeet geheel katuil dat je niet alleen komt
full caption & clip details
A young adult somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as jealousy and envy, contempt, sourness; style: monologue, storytelling; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 0.2/10; 11.9s, DUTCH.
1724_10176_000285 · in -26.5 dBFS · gain +6.5 dB · mls-00085
(anger, triumph, pride · subdued, steady, no disfluency, didactic) t is waar ook zeide hoogenberg mijne heeren mijn neef de heer bleek uit amsterdam een aktie een aktie klonk het als uit één mond hij noemt ons mijne heeren
full caption & clip details
A child somewhat feminine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; crisply articulate, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger, triumph, pride; style: didactic, whispered; good recording, quiet background; genuineness 0.2/6; vocal-burst blend 0.4/10; 14.2s, DUTCH.
1724_10176_000486 · in -25.0 dBFS · gain +5.0 dB · mls-00085
(anger, triumph, pride · subdued, steady, no disfluency, didactic) t is waar ook zeide hoogenberg mijne heeren mijn neef de heer bleek uit amsterdam een aktie een aktie klonk het als uit één mond hij noemt ons mijne heeren
full caption & clip details
A child somewhat feminine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; crisply articulate, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger, triumph, pride; style: didactic, whispered; good recording, quiet background; genuineness 0.2/6; vocal-burst blend 0.4/10; 14.2s, DUTCH.
1724_10176_000486 · in -25.0 dBFS · gain +5.0 dB · mls-00085
Relief ↓  /  Sexual Lustidentity +0.00 emotion 27 %   c-mls-PXR · #2

This chain comes from the proxy rule: the same two-sided test as above, but because Sexual Lust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Sexual Lust around average — 0.56, higher than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.41.

At the same time Relief goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.22, then +0.11, then +0.07 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 60 s · polish · mls

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.907 before conversion and 0.909 after — it rose by 0.001. Neighbour-to-neighbour the worst pair went 0.904 → 0.909. (The earlier render, with segment 1 left raw, scores 0.788 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sexual Lust moved +0.407 in the original and +0.111 after conversion — 27 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Relief, -0.297 became -0.231.

Quality. Mean predicted overall quality across the segments went 3.22 → 3.39 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.907 → 0.909 +0.001identity cos neighbours 0.904 → 0.909d_b rescored +0.407 → +0.111d_a rescored -0.297 → -0.231d_a mined -0.297d_b mined 0.407min_cos_consec (site) 0.9096min_cos_anchor (site) 0.9121dataset mlslang polishspeaker 10900total 58.7schain gain +2.0 dBseam step 1.2 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, normally alert, slightly relaxed, some disfluency
(relief · measured, fairly steady, somewhat unclear, monologue) o połowie łuku i potężnych filarach gdy wtem wtem chwycił nas za ręce i wzniósł się nieco na pościeli oczy stały mu słupem twarz trupio blada stała się teraz zieloną wtem zobaczyłem dwa cienie
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief; style: monologue, whispered; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 1.3/10; 14.4s, POLISH.
10900_6473_000479 · in -30.4 dBFS · gain +10.4 dB · mls-00108
(shame, contempt, disgust · normal-paced, fairly steady, average clarity, monologue) nie dwóch ludzi trupów czy upiorów wyszli z pod bramy i posuwali się wprost ku mnie nogi zadrżały podemną zamknąłem oczy chcąc odegnać przywidzenie ale gdym je po chwili znów otworzył zobaczyłem cztery kroki przed sobą obu braci
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, contempt, disgust; style: monologue, storytelling; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 1.7/10; 16.4s, POLISH.
10900_6473_000433 · in -29.8 dBFS · gain +9.8 dB · mls-00108
(pride, thankfulness gratitude, contentment · normal-paced, fairly steady, somewhat unclear, monologue) obaj trzymając się za ręce okropni nabrzmiali skrwawieni tacy jak ich znaleźliśmy i patrzyli obaj we mnie tak strasznie znacie mnie nie jestem lękliwy nie jestem skłonny do przywidzeń ale to wam mówię oni tam stali i ze strachu zamieniłem się w bryłę lodu
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, thankfulness gratitude, contentment; style: monologue, narration; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 2.5/10; 16.1s, POLISH.
10900_6473_000697 · in -29.4 dBFS · gain +9.4 dB · mls-00108
(sexual lust, fear, thankfulness gratitude · measured, steady, somewhat unclear, monologue) nie mogłem się poruszyć odwrócić wtedy oni zaczęli mówić tak mówić a ja słyszałem ich głos choć tam nie było powietrza tak jak was tu słyszę spytali mnie naprzód po co przychodzę
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as sexual lust, fear, thankfulness gratitude; style: monologue, narration; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.4/10; 12.4s, POLISH.
10900_6473_001035 · in -30.1 dBFS · gain +10.1 dB · mls-00109
Triumph ↓  /  Sadnessidentity +0.03 emotion 4 %   c-mls-PXR · #3

This chain comes from the proxy rule: the same two-sided test as above, but because Sadness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Sadness around average — 0.43, lower than 57 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.54.

At the same time Triumph goes the other way, from 0.95 (higher than 96 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.35. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.44, then +0.10 — a plateau around step 1, where it barely moves.

The largest step is 0.44, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 67 s · portuguese · mls

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.889 before conversion and 0.914 after — it rose by 0.025. Neighbour-to-neighbour the worst pair went 0.889 → 0.924. (The earlier render, with segment 1 left raw, scores 0.810 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.537 in the original and +0.022 after conversion — 4 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Triumph, -0.351 became -0.310.

Quality. Mean predicted overall quality across the segments went 3.19 → 3.35 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.889 → 0.914 +0.025identity cos neighbours 0.889 → 0.924d_b rescored +0.537 → +0.022d_a rescored -0.351 → -0.310d_a mined -0.351d_b mined 0.537min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang portuguesespeaker 5677total 65.8schain gain +2.7 dBseam step 1.1 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, slightly relaxed, somewhat unclear, fairly narrow pitch
(triumph, relief, hope enthusiasm optimism · measured, subdued, fairly steady, monologue) era bem um sinal de fraqueza uma demonstração de inferioridade diante daqueles povos tenazes que os guardam durante séculos tornava-se preciso reagir desenvolver o culto das tradições mantê las sempre vivazes nas memórias e nos costumes
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, relief, hope enthusiasm optimism; style: monologue, whispered; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 4.3/10; 17.0s, PORTUGUESE.
5677_4807_000446 · in -24.3 dBFS · gain +4.3 dB · mls-00120
(fatigue exhaustion · normal-paced, normally alert, fairly steady, monologue) albernaz vinha contrariado contava arranjar um número bom para a festa que ia dar e escapavalhe era quase a esperança de casamento de uma das quatro filhas que se ia das quatro porque uma delas já estava garantida graças a deus
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion; style: monologue, authoritative; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 3.4/10; 15.3s, PORTUGUESE.
5677_4807_000105 · in -23.4 dBFS · gain +3.5 dB · mls-00120
(shame, bitterness, longing · measured, normally alert, fairly steady, monologue) o crepúsculo chegava e eles entraram em casa mergulhados na melancolia da hora a decepção porém demorou dias cavalcanti o noivo de ismênia informou que nas imediações morava um literato teimoso cultivador dos contos e canções populares do brasil foram a ele
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, bitterness, longing; style: monologue; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 3.3/10; 18.4s, PORTUGUESE.
5677_4807_000046 · in -23.8 dBFS · gain +3.8 dB · mls-00120
(sadness, contentment, longing · measured, subdued, steady, monologue) era um velho poeta que teve sua fama aí pelos setenta e tantos homem doce e ingênuo que se deixara esquecer em vida como poeta e agora se entretinha em publicar coleções que ninguém lia de contos canções adágios e ditados populares
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as sadness, contentment, longing; style: monologue; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 1.6/10; 15.7s, PORTUGUESE.
5677_4807_000459 · in -24.1 dBFS · gain +4.2 dB · mls-00120
Shame ↓  /  Affectionidentity −0.00 emotion 73 %   c-mls-PXR · #4

This chain comes from the proxy rule: the same two-sided test as above, but because Affection is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Affection clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.35.

At the same time Shame goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.35. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.18, then -0.18, then +0.17, then +0.19 — not a clean run: step 2 moves back the other way by 0.18 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 77 s · portuguese · mls

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.912 before conversion and 0.909 after — it fell by 0.002. Neighbour-to-neighbour the worst pair went 0.914 → 0.886. (The earlier render, with segment 1 left raw, scores 0.800 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Affection moved +0.354 in the original and +0.259 after conversion — 73 % of the delta retained, which is most of it. On the other named axis, Shame, -0.346 became -0.036.

Quality. Mean predicted overall quality across the segments went 3.22 → 3.43 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.912 → 0.909 -0.002identity cos neighbours 0.914 → 0.886d_b rescored +0.354 → +0.259d_a rescored -0.346 → -0.036d_a mined -0.346d_b mined 0.354min_cos_consec (site) 0.9281min_cos_anchor (site) 0.9225dataset mlslang portuguesespeaker 9351total 76.1schain gain +2.9 dBseam step 1.8 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a middle-aged somewhat masculine voice · balanced body, average recording, quiet background, slightly relaxed
(shame, disappointment, anger · measured, subdued, fairly steady, whispered) como sabe há muitos desgostos contra o regente se o imperador já tivesse a idade de constituição é que era bom ia-se embora o regente e o resto pois é verdade creio que sim entretanto nunca tinha pensado nisto seriamente mas as cousas são assim mesmo que acha
full caption & clip details
A middle-aged somewhat masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is slightly cool, dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as shame, disappointment, anger; style: whispered, monologue; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 3.3/10; 17.8s, PORTUGUESE.
9351_9163_000345 · in -27.2 dBFS · gain +7.2 dB · mls-00124
(disappointment · measured, subdued, fairly steady, monologue) acho que fez bem em todo o caso peco lhe segredo não diga nada a mamãe crê que ela se oponha não mas pode ser que não se alcance nada e para lhe não dar uma esperança que pode falhar e só isto era plausível a explicação
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment; style: monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 3.9/10; 17.2s, PORTUGUESE.
9351_9163_000242 · in -28.4 dBFS · gain +8.4 dB · mls-00123
(normal-paced, normally alert, fairly steady, monologue) explicação prometi lhe não dizer nada creio que falamos ainda de política e da política daqueles últimos dez anos que não era pouca nem plácida félix não tinha certamente um plano de idéias e apreciações originais
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 3.0/10; 13.4s, PORTUGUESE.
9351_9163_000142 · in -29.2 dBFS · gain +9.2 dB · mls-00123
(contentment, thankfulness gratitude · normal-paced, normally alert, steady, monologue) através das palavras dele apalpava eu as fórmulas e os juízos do círculo ou das pessoas com quem ele lidava para o fim de encetar a carreira agora a particularidade dele era a clareza e retidão de espírito precisas para só recolher do que ouvia a parte sã e justa
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as contentment, thankfulness gratitude; style: monologue, narration; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 3.8/10; 17.8s, PORTUGUESE.
9351_9163_000023 · in -27.5 dBFS · gain +7.5 dB · mls-00123
(affection · normal-paced, normally alert, steady, monologue) ou pelo menos a porção moderada nunca andaria nos extremos qualquer que fosse o seu partido trabalhou muito hoje perguntou-me ele quando nos preparávamos para jantar
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as affection; style: monologue; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.0/10; 10.6s, PORTUGUESE.
9351_9163_000151 · in -30.2 dBFS · gain +10.2 dB · mls-00123
Contempt ↓  /  Jealousy and Envyidentity −0.00 emotion 50 %   c-mls-PXR · #5

This chain comes from the proxy rule: the same two-sided test as above, but because Jealousy and Envy is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Jealousy and Envy clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.28.

At the same time Contempt goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.44. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.08, then -0.02, then +0.07, then +0.15 — not a clean run: step 2 moves back the other way by 0.02 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.97 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 82 s · spanish · mls

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.942 before conversion and 0.940 after — it fell by 0.003. Neighbour-to-neighbour the worst pair went 0.953 → 0.927. (The earlier render, with segment 1 left raw, scores 0.837 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Jealousy and Envy moved +0.278 in the original and +0.139 after conversion — 50 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Contempt, -0.438 became -0.558.

Quality. Mean predicted overall quality across the segments went 3.13 → 3.31 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.942 → 0.940 -0.003identity cos neighbours 0.953 → 0.927d_b rescored +0.278 → +0.139d_a rescored -0.438 → -0.558d_a mined -0.438d_b mined 0.278min_cos_consec (site) 0.9716min_cos_anchor (site) 0.9618dataset mlslang spanishspeaker 3946total 80.3schain gain -0.6 dBseam step 0.3 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, steady, fairly narrow pitch
(contempt, relief, malevolence malice · normal-paced, no disfluency, clear, monologue) la respuesta que les di los primeros pueblos fue que les halagó y dixo que iria presto les ayudar y que entretanto que iba que se ayudasen de otros pueblos sus vecinos y que esperasen en campo los mexicanos
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, relief, malevolence malice; style: monologue, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.2/10; 14.5s, SPANISH.
3946_11219_005788 · in -27.9 dBFS · gain +7.9 dB · mls-00032
(sourness, affection, pain · normal-paced, no disfluency, clear, monologue) y que todos juntos les diesen guerra é que si los mexicanos viesen que les mostraban cara y ponian fuerzas contra ellos que temerían é que ya no tenian tantos poderes los mexicanos para les dar guerra como solian porque tenian muchos contrarios
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, affection, pain; style: monologue, narration; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 2.0/10; 16.9s, SPANISH.
3946_11219_004551 · in -27.0 dBFS · gain +7.0 dB · mls-00032
(shame, sadness, malevolence malice · normal-paced, no disfluency, clear, narration) y tantas palabras les dixo con nuestras lenguas tá les esforzó que reposron algo sus corazones y no tanto que luego demandron cartas para dos pueblos sus comarcanos nuestros amigos para que les fuesen ayudar
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, sadness, malevolence malice; style: narration, monologue; good recording, quiet background; genuineness 0.4/6; vocal-burst blend 1.0/10; 15.4s, SPANISH.
3946_11219_004296 · in -27.4 dBFS · gain +7.3 dB · mls-00032
(disappointment, bitterness, sadness · normal-paced, almost no disfluency, somewhat unclear, monologue) las cartas en aquel tiempo no las entendian mas bien sabian que entre nosotros se tenia por cosa cierta que guando se enviaban eran como mandamientos señales que les mandaban algunas cosas de calidad é con ellas se fueron muy contentos y las mostrron sus amigos y los llarnron
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, bitterness, sadness; style: monologue, narration; average recording, quiet background; genuineness 0.4/6; vocal-burst blend 3.0/10; 19.8s, SPANISH.
3946_11219_005327 · in -27.2 dBFS · gain +7.2 dB · mls-00032
(jealousy and envy, shame, thankfulness gratitude · measured, no disfluency, clear, monologue) y como nuestro cortes se lo mandó aguardron en el campo los mexicanos y tuvieron con ellos una batalla y con ayuda de nuestros amigos sus vecinos a quien dieron la carta no les fue mal en la pelea
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, shame, thankfulness gratitude; style: monologue, narration; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 1.7/10; 14.6s, SPANISH.
3946_11219_004879 · in -27.8 dBFS · gain +7.8 dB · mls-00032
Fear ↓  /  Jealousy and Envyidentity −0.01 emotion 75 %   c-mls-PXR · #6

This chain comes from the proxy rule: the same two-sided test as above, but because Jealousy and Envy is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Jealousy and Envy clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.28.

At the same time Fear goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.07 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 45 s · german · mls

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.896 before conversion and 0.889 after — it fell by 0.007. Neighbour-to-neighbour the worst pair went 0.902 → 0.909. (The earlier render, with segment 1 left raw, scores 0.744 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Jealousy and Envy moved +0.279 in the original and +0.209 after conversion — 75 % of the delta retained, which is most of it. On the other named axis, Fear, -0.287 became -0.629.

Quality. Mean predicted overall quality across the segments went 3.01 → 3.36 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.896 → 0.889 -0.007identity cos neighbours 0.902 → 0.909d_b rescored +0.279 → +0.209d_a rescored -0.287 → -0.629d_a mined -0.287d_b mined 0.279min_cos_consec (site) 0.9281min_cos_anchor (site) 0.9121dataset mlslang germanspeaker 252total 44.4schain gain +3.3 dBseam step 1.4 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, balanced body, no background noise, subdued, slightly relaxed, almost no disfluency, clear
(malevolence malice, fear, sadness · slow, steady, fairly narrow pitch, narration) und der alte schwarzkünstler schien dem nichts troz bieten zu wollen so sah er aus als er den teufel bannte sagte die wahrsagerin
full caption & clip details
A middle-aged masculine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, fear, sadness; style: narration, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.6/10; 11.5s, GERMAN.
252_1552_000553 · in -24.4 dBFS · gain +4.4 dB · mls-00016
(malevolence malice, bitterness, contempt · slow, fairly steady, fairly narrow pitch, storytelling) nur haben sie ihm nachher die hände gefaltet daß er hier unten wider willen beten muß und warum betet er denn fragte ich zornig da drüben über uns im himmelssee funkeln und schwimmen zwar unzählige sterne
full caption & clip details
A middle-aged masculine voice; delivery is subdued, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, rough, balanced body; clear, almost no disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, bitterness, contempt; style: storytelling, narration; average recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.1/10; 15.4s, GERMAN.
252_1552_000588 · in -21.0 dBFS · gain +1.0 dB · mls-00016
(jealousy and envy, pride, contempt · measured, fairly steady, moderate pitch range, storytelling) aber wenn es welten sind wie viele kluge köpfe behaupten so giebt es auch schädel auf ihnen und würmer wie hier unten das geht so fort durch die ganze unermeßlichkeit und der baseler todtentanz wird dadurch nur um so lustiger und wilder und der ballsaal größer
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, pride, contempt; style: storytelling, monologue; average recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.0/10; 18.0s, GERMAN.
252_1552_000213 · in -21.2 dBFS · gain +1.2 dB · mls-00016
Emotional Numbness ↓  /  Sournessidentity −0.05 emotion 55 %   c-mls-PXR · #7

This chain comes from the proxy rule: the same two-sided test as above, but because Sourness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Sourness clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.27.

At the same time Emotional Numbness goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.07, then +0.21 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 47 s · spanish · mls

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.934 before conversion and 0.882 after — it fell by 0.052. Neighbour-to-neighbour the worst pair went 0.942 → 0.889. (The earlier render, with segment 1 left raw, scores 0.832 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sourness moved +0.275 in the original and +0.152 after conversion — 55 % of the delta retained. On the other named axis, Emotional Numbness, -0.254 became -0.348.

Quality. Mean predicted overall quality across the segments went 3.22 → 3.30 (+0.08) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.934 → 0.882 -0.052identity cos neighbours 0.942 → 0.889d_b rescored +0.275 → +0.152d_a rescored -0.254 → -0.348d_a mined -0.254d_b mined 0.275min_cos_consec (site) 0.9458min_cos_anchor (site) 0.9383dataset mlslang spanishspeaker 9972total 46.2schain gain +2.3 dBseam step 0.4 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(emotional numbness, malevolence malice, pride · steady, no disfluency, monologue, authoritative) y volvió á salir á la mar y toda la gente venía á él y los enseñaba y pasando vió á leví hijo de alfeo sentado al banco de los públicos tributos y le dice
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness, malevolence malice, pride; style: monologue, authoritative; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.6/10; 12.7s, SPANISH.
9972_11159_000008 · in -31.2 dBFS · gain +11.2 dB · mls-00031
(contentment, pride, malevolence malice · fairly steady, no disfluency, monologue, authoritative) sígueme y levantándose le siguió y aconteció que estando jesús á la mesa en casa de él muchos publicanos y pecadores estaban también á la mesa juntamente con jesús y con sus discípulos porque había muchos y le habían seguido
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contentment, pride, malevolence malice; style: monologue, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.6/10; 18.5s, SPANISH.
9972_11159_000026 · in -32.1 dBFS · gain +12.1 dB · mls-00031
(sourness, disgust, jealousy and envy · fairly steady, almost no disfluency, monologue, authoritative) y los escribas y los fariseos viéndole comer con los publicanos y con los pecadores dijeron á sus discípulos qué es esto que él come y bebe con los publicanos y con los pecadores y oyéndolo jesús les dice
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, disgust, jealousy and envy; style: monologue, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.1/10; 15.4s, SPANISH.
9972_11159_000005 · in -32.4 dBFS · gain +12.4 dB · mls-00031
Relief ↓  /  Angeridentity −0.03 emotion 56 %   c-mls-PXR · #8

This chain comes from the proxy rule: the same two-sided test as above, but because Anger is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Anger around average — 0.58, higher than 58 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.42.

At the same time Relief goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.63 (higher than 63 % of clips in this corpus), a change of -0.35. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.00, then +0.19, then +0.00, then +0.23 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 66 s · dutch · mls

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.947 before conversion and 0.922 after — it fell by 0.026. Neighbour-to-neighbour the worst pair went 0.938 → 0.927. (The earlier render, with segment 1 left raw, scores 0.764 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Anger moved +0.416 in the original and +0.232 after conversion — 56 % of the delta retained. On the other named axis, Relief, -0.350 became -0.329.

Quality. Mean predicted overall quality across the segments went 3.20 → 3.52 (+0.32) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.947 → 0.922 -0.026identity cos neighbours 0.938 → 0.927d_b rescored +0.416 → +0.232d_a rescored -0.350 → -0.329d_a mined -0.350d_b mined 0.420min_cos_consec (site) 0.9418min_cos_anchor (site) 0.9533dataset mlslang dutchspeaker 1724total 64.4schain gain +2.8 dBseam step 1.0 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a child somewhat feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, measured, slightly relaxed, light breath
(relief, fatigue exhaustion · subdued, fairly steady, frequent disfluency, monologue) of vervolgde bol de septer van ageen dan iets anders ware geweest dan een stok ofschoon ik beken nergens gelezen te hebben dat hij zich met een tang behielp voegde hij er halfluid bij
full caption & clip details
A child somewhat feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as relief, fatigue exhaustion; style: monologue, narration; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.0/10; 13.5s, DUTCH.
1724_10176_000469 · in -28.5 dBFS · gain +8.5 dB · mls-00085
(relief, fatigue exhaustion · subdued, fairly steady, frequent disfluency, monologue) of vervolgde bol de septer van ageen dan iets anders ware geweest dan een stok ofschoon ik beken nergens gelezen te hebben dat hij zich met een tang behielp voegde hij er halfluid bij
full caption & clip details
A child somewhat feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as relief, fatigue exhaustion; style: monologue, narration; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.0/10; 13.5s, DUTCH.
1724_10176_000469 · in -28.5 dBFS · gain +8.5 dB · mls-00085
(jealousy and envy, contempt, sourness · normally alert, fairly steady, some disfluency, monologue) maar hervatte hij zich herinnerende welken plicht hij als gastheer te vervullen had wij zullen dat punt te zijner tijd behandelen je vergeet geheel katuil dat je niet alleen komt
full caption & clip details
A young adult somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as jealousy and envy, contempt, sourness; style: monologue, storytelling; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 0.2/10; 11.9s, DUTCH.
1724_10176_000285 · in -26.5 dBFS · gain +6.5 dB · mls-00085
(jealousy and envy, contempt, sourness · normally alert, fairly steady, some disfluency, monologue) maar hervatte hij zich herinnerende welken plicht hij als gastheer te vervullen had wij zullen dat punt te zijner tijd behandelen je vergeet geheel katuil dat je niet alleen komt
full caption & clip details
A young adult somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as jealousy and envy, contempt, sourness; style: monologue, storytelling; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 0.2/10; 11.9s, DUTCH.
1724_10176_000285 · in -26.5 dBFS · gain +6.5 dB · mls-00085
(anger, triumph, pride · subdued, steady, no disfluency, didactic) t is waar ook zeide hoogenberg mijne heeren mijn neef de heer bleek uit amsterdam een aktie een aktie klonk het als uit één mond hij noemt ons mijne heeren
full caption & clip details
A child somewhat feminine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; crisply articulate, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger, triumph, pride; style: didactic, whispered; good recording, quiet background; genuineness 0.2/6; vocal-burst blend 0.4/10; 14.2s, DUTCH.
1724_10176_000486 · in -25.0 dBFS · gain +5.0 dB · mls-00085
Anger ↓  /  Longingidentity +0.01 emotion 41 %   c-mls-PXR · #9

This chain comes from the proxy rule: the same two-sided test as above, but because Longing is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Longing clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.29.

At the same time Anger goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.08, then -0.07, then +0.12 — not a clean run: step 3 moves back the other way by 0.07 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 82 s · dutch · mls

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.929 before conversion and 0.943 after — it rose by 0.013. Neighbour-to-neighbour the worst pair went 0.925 → 0.943. (The earlier render, with segment 1 left raw, scores 0.725 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.295 in the original and +0.119 after conversion — 41 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Anger, -0.320 became -0.171.

Quality. Mean predicted overall quality across the segments went 3.21 → 3.40 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.929 → 0.943 +0.013identity cos neighbours 0.925 → 0.943d_b rescored +0.295 → +0.119d_a rescored -0.320 → -0.171d_a mined -0.320d_b mined 0.294min_cos_consec (site) 0.9313min_cos_anchor (site) 0.9330dataset mlslang dutchspeaker 1724total 80.5schain gain +2.6 dBseam step 1.0 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an elderly somewhat feminine voice · neutral-toned, fairly smooth, normally alert, fairly steady, clear, light breath
(anger, bitterness, thankfulness gratitude · measured, slightly relaxed, some disfluency, monologue) uitmuntend antwoordde roda en wreef zich vergenoegd de handen dan gaan we eerst eens kijken hoe het leven daar op de stations ons bevalt is het er voor ons beiden te druk dan keer ik met moeder naar brisbane terug en we blijven er wonen tot we samen naar nederland terugkeeren
full caption & clip details
An elderly somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger, bitterness, thankfulness gratitude; style: monologue, narration; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.0/10; 17.3s, DUTCH.
1724_2649_000174 · in -24.5 dBFS · gain +4.5 dB · mls-00103
(bitterness, shame · measured, slightly relaxed, almost no disfluency, narration) de geheele familie zat onder de veranda van darlingstation vereenigd willem vertelde van zijn laatste bezoek op den kruisberg en bij jan kranse en toch willem wed ik dat je in nederland nog iets vergeten hebt wat je beloofd hebt te doen zeide knol opeens ik wed van niet antwoordde willem wat bedoel je kees
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as bitterness, shame; style: narration, whispered; average recording, quiet background; genuineness 0.8/6; vocal-burst blend 1.3/10; 19.4s, DUTCH.
1724_2649_000943 · in -25.9 dBFS · gain +5.9 dB · mls-00103
(astonishment surprise, hope enthusiasm optimism, affection · measured, neutral tension, almost no disfluency, narration) wel je hebt de jongens van het gymnasium in doetinchem meer dan eens beloofd te schrijven hoe het ons op onze vlucht gegaan is t is niet mooi dat je het vergeten hebt willem zonder hen zouden we nooit vr onzen tijd van den kruisberg zijn afgekomen
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as astonishment surprise, hope enthusiasm optimism, affection; style: narration, whispered; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 2.5/10; 14.6s, DUTCH.
1724_2649_000670 · in -25.0 dBFS · gain +5.0 dB · mls-00103
(disappointment, affection, doubt · normal-paced, neutral tension, some disfluency, whispered) ik heb het niet vergeten kees tot mijne spijt moet ik echter bekennen dat ik hunne naamkaartjes heb verloren maar ik weet goeden raad niemand zal me kunnen verwijten dat ik een belofte heb geschonden zoo spoedig ik tijd heb zal ik mijne geschiedenis eens opschrijven en laten drukken
full caption & clip details
A child somewhat masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; clear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as disappointment, affection, doubt; style: whispered, monologue; good recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.8/10; 16.2s, DUTCH.
1724_2649_001908 · in -25.1 dBFS · gain +5.1 dB · mls-00103
(longing, awe, emotional numbness · measured, slightly relaxed, almost no disfluency, narration) wellicht krijgt een van die jongens het boek in handen en dan hebben ze voor het lange wachten mijne overige lotgevallen op den koop toe eli heimans
full caption & clip details
An elderly somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing, awe, emotional numbness; style: narration, whispered; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 1.3/10; 13.7s, DUTCH.
1724_2649_000091 · in -24.9 dBFS · gain +4.9 dB · mls-00103
Contemplation ↓  /  Thankfulness Gratitudeidentity −0.01 emotion 94 %   c-mls-PXR · #10

This chain comes from the proxy rule: the same two-sided test as above, but because Thankfulness Gratitude is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Thankfulness Gratitude around average — 0.57, higher than 57 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.42.

At the same time Contemplation goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.17, then +0.20, then +0.05 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 66 s · dutch · mls

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.943 before conversion and 0.938 after — it fell by 0.006. Neighbour-to-neighbour the worst pair went 0.905 → 0.920. (The earlier render, with segment 1 left raw, scores 0.760 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.418 in the original and +0.394 after conversion — 94 % of the delta retained, which is essentially all of it. On the other named axis, Contemplation, -0.333 became -0.449.

Quality. Mean predicted overall quality across the segments went 3.18 → 3.40 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.943 → 0.938 -0.006identity cos neighbours 0.905 → 0.920d_b rescored +0.418 → +0.394d_a rescored -0.333 → -0.449d_a mined -0.333d_b mined 0.418min_cos_consec (site) 0.9418min_cos_anchor (site) 0.9555dataset mlslang dutchspeaker 1724total 64.7schain gain +1.6 dBseam step 1.2 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult feminine voice · neutral-toned, neutral-bright, quiet background, slightly relaxed, fairly steady, clear
(contemplation, confusion, contentment · normal-paced, normally alert, some disfluency, narration) myn vrouw was konfuus en ik dacht er aan wat busselinck en waterman zouden gezegd hebben als ze dat gezien hadden dat er twee rytuigen tegelyk voor ons waren meen ik maar t was niet gemakkelyk een keus te doen want ik kon niet besluiten een der partyen te krenken door t afwyzen van een zoo lieve attentie
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as contemplation, confusion, contentment; style: narration, whispered; good recording, quiet background; genuineness 1.3/6; vocal-burst blend 1.5/10; 15.9s, DUTCH.
1724_1362_001141 · in -25.9 dBFS · gain +5.8 dB · mls-00094
(jealousy and envy, pride, contemplation · measured, normally alert, some disfluency, narration) goede raad was duur maar ik heb my uit die hoogstmoeielyke omstandigheid alweer gered ik heb myn vrouw en marie in t roode rytuig gezet in den wagen van t rooie vest meen ik en ik ben in t gele gaan zitten in t gele rytuig meen ik
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, pride, contemplation; style: narration, whispered; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 0.0/10; 14.3s, DUTCH.
1724_1362_002135 · in -26.8 dBFS · gain +6.8 dB · mls-00094
(bitterness, awe, anger · slow, very low-energy, some disfluency, narration) wat die paarden liepen op de weesperstraat waar t altyd zoo vuil is vloog de modder rechts en links huizenhoog en alsof weer t spel sprak daar liep die schooierige sjaalman in gebogen houding met gebukt hoofd en ik zag hoe hy met de mouw van zyn kaal jasje zyn bleek gelaat trachtte te reinigen van de spatten
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, thin; clear, some disfluency, fairly narrow pitch, audible breath; affect is mildly positive, neutral stance, slightly guarded; reads as bitterness, awe, anger; style: narration, storytelling; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 1.5/10; 19.3s, DUTCH.
1724_1362_001214 · in -26.6 dBFS · gain +6.6 dB · mls-00094
(thankfulness gratitude, sourness, jealousy and envy · normal-paced, normally alert, no disfluency, narration) wat het in den tekst behandeld voorval aangaat de officier van gezondheid bensen heeft kort na t verschynen van den havelaar in de n rotterdamsche courant meegedeeld dat de heer carolus na z'n tehuiskomst van parang koedjang niet weinige uren had geleefd maar nog ik meen twee dagen
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as thankfulness gratitude, sourness, jealousy and envy; style: narration, storytelling; average recording, quiet background; genuineness 0.6/6; vocal-burst blend 0.0/10; 15.7s, DUTCH.
1724_1362_000041 · in -26.4 dBFS · gain +6.4 dB · mls-00093
Emotional Numbness ↓  /  Triumphidentity −0.05 emotion 48 %   c-mls-PXR · #11

This chain comes from the proxy rule: the same two-sided test as above, but because Triumph is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Triumph clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.34.

At the same time Emotional Numbness goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.62 (higher than 62 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.20 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 39 s · italian · mls

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.917 before conversion and 0.866 after — it fell by 0.051. Neighbour-to-neighbour the worst pair went 0.924 → 0.832. (The earlier render, with segment 1 left raw, scores 0.802 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Triumph moved +0.341 in the original and +0.165 after conversion — 48 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Emotional Numbness, -0.375 became -0.268.

Quality. Mean predicted overall quality across the segments went 3.22 → 3.33 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.917 → 0.866 -0.051identity cos neighbours 0.924 → 0.832d_b rescored +0.341 → +0.165d_a rescored -0.375 → -0.268d_a mined -0.375d_b mined 0.341min_cos_consec (site) 0.9198min_cos_anchor (site) 0.9164dataset mlslang italianspeaker 2019total 38.3schain gain +2.2 dBseam step 0.6 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, average recording, quiet background, subdued, slightly relaxed, steady, frequent disfluency, fairly narrow pitch
(emotional numbness, contentment, contemplation · measured, somewhat unclear, light breath, monologue) del vizio egli fa sentire il puzzo non il profumo le sue nudità son nudità di tavola anatomica che non ispirano il menomo pensiero sensuale non c'è nessuno dei suoi libri neanche il più crudo
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, contentment, contemplation; style: monologue, whispered; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 3.0/10; 16.6s, ITALIAN.
2019_1577_001291 · in -25.1 dBFS · gain +5.1 dB · mls-00068
(awe, infatuation, pain · slow, slurred, minimal breath, narration) che non lasci nell'animo netta ferma immutabile l'avversione o il disprezzo per le basse passioni che vi sono trattate
full caption & clip details
A child somewhat feminine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; slurred, frequent disfluency, fairly narrow pitch, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as awe, infatuation, pain; style: narration, monologue; average recording, quiet background; genuineness 0.1/6; vocal-burst blend 0.0/10; 10.1s, ITALIAN.
2019_1577_000625 · in -22.8 dBFS · gain +2.8 dB · mls-00067
(triumph, pride, thankfulness gratitude · slow, slurred, audible breath, narration) egli non è come il dumas figlio legato da un'invincibile simpatia ai suoi mostri di donne a cui dice infami ad alta voce o care a
full caption & clip details
A child somewhat feminine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, pride, thankfulness gratitude; style: narration, monologue; average recording, quiet background; genuineness 0.3/6; vocal-burst blend 0.6/10; 12.1s, ITALIAN.
2019_1577_001039 · in -23.5 dBFS · gain +3.5 dB · mls-00068
Jealousy and Envy ↓  /  Longingidentity +0.01 emotion 56 %   c-mls-PXR · #12

This chain comes from the proxy rule: the same two-sided test as above, but because Longing is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Longing clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.29.

At the same time Jealousy and Envy goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.50 (right about the corpus median), a change of -0.46. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.14, then +0.19, then -0.04 — not a clean run: step 3 moves back the other way by 0.04 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 58 s · french · mls

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.870 before conversion and 0.877 after — it rose by 0.007. Neighbour-to-neighbour the worst pair went 0.855 → 0.852. (The earlier render, with segment 1 left raw, scores 0.617 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.288 in the original and +0.160 after conversion — 56 % of the delta retained. On the other named axis, Jealousy and Envy, -0.461 became -0.586.

Quality. Mean predicted overall quality across the segments went 2.65 → 3.23 (+0.58) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.870 → 0.877 +0.007identity cos neighbours 0.855 → 0.852d_b rescored +0.288 → +0.160d_a rescored -0.461 → -0.586d_a mined -0.461d_b mined 0.288min_cos_consec (site) 0.8988min_cos_anchor (site) 0.9249dataset mlslang frenchspeaker 12899total 57.0schain gain +2.9 dBseam step 0.4 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · slightly cool, neutral-bright, fairly smooth, thin, average recording, normally alert, slightly relaxed, some disfluency
(jealousy and envy, contentment, relief · normal-paced, fairly steady, clear, didactic) la petite jane avait trois ans quand elle perdit sa mère elle devint la consolation de sa grand'mère et de sa tante et tout semblait présager qu'elle était fixée à highbury pour la vie mais l'intervention d'un ami de son père modifia sa destinée
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, contentment, relief; style: didactic, monologue; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 1.5/10; 17.6s, FRENCH.
12899_10625_001632 · in -28.2 dBFS · gain +8.2 dB · mls-00045
(awe, jealousy and envy · normal-paced, moderately variable, average clarity, playful) le colonel campbell tenait en grande estime le lieutenant fairfax et de plus il considérait devoir la vie aux soins dont son compagnon d'armes l'avait entouré pendant les accès d'une fièvre contractée au cours d'une campagne il demeura fidèle à la mémoire de son ami
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as awe, jealousy and envy; style: playful, didactic; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 4.0/10; 18.5s, FRENCH.
12899_10625_001490 · in -27.9 dBFS · gain +7.9 dB · mls-00045
(sadness, disappointment, distress · measured, moderately variable, average clarity, monologue) et bien que plusieurs années se fussent écoulées entre la mort du pauvre fairfax et le retour du colonel en angleterre sa reconnaissance n'en fut pas affaiblie
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sadness, disappointment, distress; style: monologue, authoritative; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 1.9/10; 11.2s, FRENCH.
12899_10625_001851 · in -28.4 dBFS · gain +8.4 dB · mls-00046
(longing, malevolence malice · measured, moderately variable, clear, monologue) dès son arrivée il s'occupa de rechercher l'enfant et s'intéressa à elle le colonel était marié et avait une fille à peu près de l'âge de jane
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as longing, malevolence malice; style: monologue, whispered; average recording, no background noise; genuineness 1.9/6; vocal-burst blend 0.9/10; 10.2s, FRENCH.
12899_10625_002026 · in -28.3 dBFS · gain +8.3 dB · mls-00046
Contemplation ↓  /  Reliefidentity −0.05 emotion 67 %   c-mls-PXR · #13

This chain comes from the proxy rule: the same two-sided test as above, but because Relief is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Relief clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.31.

At the same time Contemplation goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.57 (higher than 57 % of clips in this corpus), a change of -0.39. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.16 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 50 s · german · mls

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.957 before conversion and 0.906 after — it fell by 0.052. Neighbour-to-neighbour the worst pair went 0.950 → 0.872. (The earlier render, with segment 1 left raw, scores 0.795 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.313 in the original and +0.210 after conversion — 67 % of the delta retained. On the other named axis, Contemplation, -0.395 became -0.412.

Quality. Mean predicted overall quality across the segments went 3.06 → 3.26 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.957 → 0.906 -0.052identity cos neighbours 0.950 → 0.872d_b rescored +0.313 → +0.210d_a rescored -0.395 → -0.412d_a mined -0.394d_b mined 0.313min_cos_consec (site) 0.9562min_cos_anchor (site) 0.9647dataset mlslang germanspeaker 4174total 49.6schain gain +4.0 dBseam step 0.4 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normal-paced, normally alert, slightly relaxed
(contemplation, concentration, contentment · steady, formal, monologue) die die große mehrheit des volkes von den angelegenheiten des staates ausschlossen nur wenn wir das überwinden werden wir zu einem deutschen gesamtvolk kommen dem auch wieder die achtung der welt zuteil werden wird
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contemplation, concentration, contentment; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 15.0s, GERMAN.
4174_10388_000006 · in -27.0 dBFS · gain +7.0 dB · mls-00003
(disappointment, emotional numbness, concentration · fairly steady, formal, monologue) diese ausführungen sagten einem teil der anwesenden wenig zu der gast aus salzburg wurde lärmend unterbrochen es läßt sich nachfühlen daß den herren ein wiener faszist der die reigen aufführungen mit stinkbomben heimsuchte lieber gewesen wäre als dieser ehrliche demokrat
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, emotional numbness, concentration; style: formal, monologue; good recording, quiet background; genuineness 0.1/6; vocal-burst blend 0.0/10; 19.4s, GERMAN.
4174_10388_000114 · in -27.1 dBFS · gain +7.1 dB · mls-00003
(relief · fairly steady, monologue, formal) die diskussion war langatmig und ohne niveau abgeordneter heile und stefan großmann versuchten vergeblich gedanken hineinzutragen graf montgelas überbrachte dem salzburger einen gruß des benachbarten
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief; style: monologue, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 15.6s, GERMAN.
4174_10388_000043 · in -27.0 dBFS · gain +7.0 dB · mls-00003
Helplessness ↓  /  Angeridentity +0.07 emotion 64 %   c-mls-PXR · #14

This chain comes from the proxy rule: the same two-sided test as above, but because Anger is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Anger clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.30.

At the same time Helplessness goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.41 (lower than 59 % of clips in this corpus), a change of -0.57. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.06, then +0.08, then -0.01 — not a clean run: step 4 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 81 s · portuguese · mls

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.829 before conversion and 0.903 after — it rose by 0.074. Neighbour-to-neighbour the worst pair went 0.864 → 0.901. (The earlier render, with segment 1 left raw, scores 0.738 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Anger moved +0.298 in the original and +0.191 after conversion — 64 % of the delta retained. On the other named axis, Helplessness, -0.215 became -0.532.

Quality. Mean predicted overall quality across the segments went 3.27 → 3.43 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.829 → 0.903 +0.074identity cos neighbours 0.864 → 0.901d_b rescored +0.298 → +0.191d_a rescored -0.215 → -0.532d_a mined -0.567d_b mined 0.297min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang portuguesespeaker 10107total 80.0schain gain +1.6 dBseam step 0.8 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an elderly masculine voice
(shame, helplessness, malevolence malice · measured, normally alert, slightly relaxed, narration) mancebos fracos e degenerados de hoje sois incapazes de encurvar o arco de vossos antepassados em outros tempos quando a idade nâo tinha ainda branqueado estes cabellos nem quebrado estes pulsos
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is warm, neutral-bright, slightly rough, full; very clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as shame, helplessness, malevolence malice; style: narration, storytelling; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 1.1/10; 15.0s, PORTUGUESE.
10107_6390_000076 · in -31.1 dBFS · gain +11.1 dB · mls-00121
(disgust, pride, awe · slow, normally alert, slightly relaxed, narration) eu só ou qualquer dos meus valentes teria esmagado este mancebo com a mesma facilidade com que espedaço este cachimbo e esmagou entre os dedos o canudo pelo qual aspirava a fumaça da pituma
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, balanced body; slurred, no disfluency, very wide pitch range, light breath; affect is positive, neutral stance, slightly guarded; reads as disgust, pride, awe; style: narration, storytelling; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.3/10; 15.5s, PORTUGUESE.
10107_6390_000021 · in -30.1 dBFS · gain +10.1 dB · mls-00121
(disgust, infatuation, awe · measured, very low-energy, slightly relaxed, narration) a este gesto a estas dnnis palavras bagas de suor frio escorregaram ela testa do joven guerreiro que batendo os dentes como um queixada enfurecido com voz convulsa e abafada respondeu
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; very clear, almost no disfluency, very wide pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, infatuation, awe; style: narration, storytelling; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.5/10; 15.4s, PORTUGUESE.
10107_6390_000177 · in -30.6 dBFS · gain +10.6 dB · mls-00121
(sourness, anger, contempt · measured, energised, slightly relaxed, storytelling) oriçanga oriçanga não profiras taes palavras a cólera te cega velho cacitjue e lorna le injusto nào penses que esse estrangeiro que acabámos de garrotear era um inimigo vulgar nào era ura enviado de anhangá
full caption & clip details
An elderly masculine voice; delivery is energised, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; very clear, frequent disfluency, very wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as sourness, anger, contempt; style: storytelling, monologue; average recording, no background noise; genuineness 1.7/6; vocal-burst blend 1.4/10; 16.9s, PORTUGUESE.
10107_6390_000160 · in -29.9 dBFS · gain +9.9 dB · mls-00121
(anger, disgust, fear · normal-paced, energised, neutral tension, storytelling) e estou certo que cora elle combatiam contra nós os manitós das trevas occultos entre os ramos da floresta se lá te acharas se presenciasses esse estranho combate e visses por que modo sobrenatural o maldito emboaba se furtava a nossos golpes
full caption & clip details
A child masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as anger, disgust, fear; style: storytelling, cartoonish; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 4.6/10; 17.9s, PORTUGUESE.
10107_6390_000178 · in -29.1 dBFS · gain +9.2 dB · mls-00121
Emotional Numbness ↓  /  Contentmentidentity −0.04 emotion 132 %   c-mls-PXR · #15

This chain comes from the proxy rule: the same two-sided test as above, but because Contentment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contentment clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.29.

At the same time Emotional Numbness goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.49 (right about the corpus median), a change of -0.47. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.11, then +0.18 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 44 s · dutch · mls

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.955 before conversion and 0.914 after — it fell by 0.041. Neighbour-to-neighbour the worst pair went 0.945 → 0.908. (The earlier render, with segment 1 left raw, scores 0.796 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.289 in the original and +0.383 after conversion — 132 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, -0.475 became -0.171.

Quality. Mean predicted overall quality across the segments went 3.23 → 3.43 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.955 → 0.914 -0.041identity cos neighbours 0.945 → 0.908d_b rescored +0.289 → +0.383d_a rescored -0.475 → -0.171d_a mined -0.475d_b mined 0.290min_cos_consec (site) 0.9650min_cos_anchor (site) 0.9650dataset mlslang dutchspeaker 1724total 43.7schain gain +1.6 dBseam step 0.6 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · quiet background
(emotional numbness, astonishment surprise · normal-paced, normally alert, slightly relaxed, monologue) dit antwoord der bekoorlijke zwerfster bevestigde een vermoeden hetwelk ik reeds den vorigen avond uit sommige van haar uitdrukkingen omtrent haar geloofsbelijdenis had opgevat ik verlangde echter geen theologische wending aan ons gesprek te geven en ving dus aan over andere onderwerpen te spreken
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, astonishment surprise; style: monologue, narration; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.0/10; 16.4s, DUTCH.
1724_1601_001708 · in -31.3 dBFS · gain +11.3 dB · mls-00096
(sourness, fatigue exhaustion, longing · normal-paced, normally alert, slightly relaxed, narration) daar ik mij echter weinig meer herinner van hetgeen wij verder afhandelden en zulks ook voor mijn lezers waarschijnlijk even vervelend zoude zijn als indien ik hen dwong de reis in de trekschuit zelve te maken zal ik slechts datgene vermelden hetwelk amelia mij toevoegde toen wij amsterdam bijna bereikt hadden
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, fatigue exhaustion, longing; style: narration, monologue; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 2.2/10; 16.4s, DUTCH.
1724_1601_001916 · in -30.6 dBFS · gain +10.6 dB · mls-00096
(contentment · slow, very low-energy, relaxed, whispered) hier zeide zij op een plechtigen toon moet onze korte kennis eindigen zoodra wij uit het oog des schippers zijn verlaten wij elkander waarschijnlijk voor altijd
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, submissive, slightly guarded; reads as contentment; style: whispered, monologue; poor recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.0/10; 11.3s, DUTCH.
1724_1601_001657 · in -34.0 dBFS · gain +14.0 dB · mls-00096
Contempt ↓  /  Concentrationidentity +0.02 emotion 50 %   c-mls-PXR · #16

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.29.

At the same time Contempt goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.10 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 39 s · dutch · mls

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.862 before conversion and 0.878 after — it rose by 0.016. Neighbour-to-neighbour the worst pair went 0.867 → 0.882. (The earlier render, with segment 1 left raw, scores 0.657 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.289 in the original and +0.144 after conversion — 50 % of the delta retained. On the other named axis, Contempt, -0.337 became +0.266.

Quality. Mean predicted overall quality across the segments went 3.04 → 3.39 (+0.36) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.862 → 0.878 +0.016identity cos neighbours 0.867 → 0.882d_b rescored +0.289 → +0.144d_a rescored -0.337 → +0.266d_a mined -0.338d_b mined 0.290min_cos_consec (site) 0.8939min_cos_anchor (site) 0.8939dataset mlslang dutchspeaker 2450total 38.6schain gain +0.4 dBseam step 0.6 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an elderly masculine voice · neutral-toned, slightly dark, average recording, slightly relaxed, fairly narrow pitch, audible breath
(contempt, impatience and irritability, anger · slow, very low-energy, fairly steady, narration) innen kort misschien zou hij weder ontmoeten treelende was aan de eene zij die gedachte maar hoedanig zou hare ontmoeting zijn
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, balanced body; average clarity, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, impatience and irritability, anger; style: narration, storytelling; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.0/10; 10.6s, DUTCH.
2450_10646_001379 · in -27.5 dBFS · gain +7.5 dB · mls-00088
(sexual lust · measured, normally alert, steady, narration) daar zij hem zonder haren broeder zag terugkeeren n wie kon hem verzekeren dat eene afwezendheid van zoovele maanden geene verandering in hare denkbeelden ten zijnen aanzien had te weeg
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; clear, almost no disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as sexual lust; style: narration, storytelling; average recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.1/10; 14.0s, DUTCH.
2450_10646_001843 · in -29.1 dBFS · gain +9.1 dB · mls-00088
(concentration, confusion, contemplation · measured, normally alert, fairly steady, storytelling) door en tot derzelver beschouwing opgewekt gedurig terug tot zijne bepeinzing en ontwaakte daaruit niet voor dat hij op den dijk gekomen de stad otterdam zelve in het oog kreeg
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, thin; slurred, some disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, confusion, contemplation; style: storytelling, narration; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 1.4/10; 14.4s, DUTCH.
2450_10646_001697 · in -27.9 dBFS · gain +7.9 dB · mls-00088
Anger ↓  /  Concentrationidentity −0.00 emotion 197 %   c-mls-PXR · #17

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.30.

At the same time Anger goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.56 (higher than 56 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.14 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.97 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 44 s · dutch · mls

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.916 before conversion and 0.913 after — it fell by 0.003. Neighbour-to-neighbour the worst pair went 0.916 → 0.913. (The earlier render, with segment 1 left raw, scores 0.712 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.303 in the original and +0.599 after conversion — 197 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Anger, -0.406 became -0.379.

Quality. Mean predicted overall quality across the segments went 3.16 → 3.39 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.916 → 0.913 -0.003identity cos neighbours 0.916 → 0.913d_b rescored +0.303 → +0.599d_a rescored -0.406 → -0.379d_a mined -0.406d_b mined 0.304min_cos_consec (site) 0.9659min_cos_anchor (site) 0.9631dataset mlslang dutchspeaker 1724total 43.1schain gain +2.6 dBseam step 0.8 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child somewhat feminine voice · average recording, quiet background, normally alert, slightly relaxed
(anger, doubt, relief · normal-paced, moderately variable, some disfluency, storytelling) waart dat's mijne privilege machteld is mijn zusterke strijdt het dan niet met de courtoisie zijne zuster te kwellen vroeg jacob jansz gij den zoon van schepen meerman courtoisie leeren riep de twaalfjarige knaap met kluchtige fierheid mij die mijne opvoeding heb ontvangen tegelijk met de jonkvrouw van egmond
full caption & clip details
A child somewhat feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as anger, doubt, relief; style: storytelling, narration; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 2.1/10; 19.4s, DUTCH.
1724_2757_001018 · in -23.3 dBFS · gain +3.3 dB · mls-00104
(fear, awe · measured, fairly steady, frequent disfluency, whispered) allen schaterden van lachen over de aanmatiging van den knaap hoort den kleuter riep guurt groot de hand nog warm van s meesters plak hoort hem praten of hij de roede al ontwassen ware
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, awe; style: whispered, monologue; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.0/10; 11.7s, DUTCH.
1724_2757_001404 · in -22.7 dBFS · gain +2.7 dB · mls-00104
(concentration, intoxication altered states of consciousness · measured, moderately variable, frequent disfluency, whispered) me maar guurt sar me maar sinds je roeit laat ik je met vrede maar als we straks aan land zijn zal ik het je betaald zetten liefelijk gerritmaat sprak jacob jansz zich tot hem omkeerend
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is slightly cool, slightly dark, fairly smooth, slightly thin; clear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, intoxication altered states of consciousness; style: whispered, monologue; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 0.8/10; 12.4s, DUTCH.
1724_2757_001201 · in -22.3 dBFS · gain +2.3 dB · mls-00104
Longing ↓  /  Shameidentity −0.02 emotion 128 %   c-mls-PXR · #18

This chain comes from the proxy rule: the same two-sided test as above, but because Shame is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Shame clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.27.

At the same time Longing goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.03 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 41 s · italian · mls

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.906 before conversion and 0.886 after — it fell by 0.020. Neighbour-to-neighbour the worst pair went 0.906 → 0.886. (The earlier render, with segment 1 left raw, scores 0.777 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.272 in the original and +0.349 after conversion — 128 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Longing, -0.258 became -0.103.

Quality. Mean predicted overall quality across the segments went 3.27 → 3.37 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.906 → 0.886 -0.020identity cos neighbours 0.906 → 0.886d_b rescored +0.272 → +0.349d_a rescored -0.258 → -0.103d_a mined -0.258d_b mined 0.272min_cos_consec (site) 0.9233min_cos_anchor (site) 0.9233dataset mlslang italianspeaker 6807total 40.0schain gain +3.7 dBseam step 0.3 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, clear, light breath
(longing, distress, pain · normal-paced, slightly relaxed, fairly steady, monologue) la provvidenza l'avevano rimorchiata a riva tutta sconquassata così come l'avevano trovata di là dal capo dei mulini col naso fra gli scogli e la schiena in aria
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as longing, distress, pain; style: monologue, authoritative; good recording, quiet background; genuineness 1.0/6; vocal-burst blend 1.0/10; 10.3s, ITALIAN.
6807_7050_000345 · in -26.8 dBFS · gain +6.8 dB · mls-00078
(contentment, shame, relief · brisk, neutral tension, moderately variable, cartoonish) in un momento era corso sulla riva tutto il paese uomini e donne e padron ntoni mischiato nella folla guardava anche lui come gli altri curiosi alcuni davano pure un calcio nella pancia della provvidenza per far suonare com'era fessa quasi non fosse più di nessuno
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contentment, shame, relief; style: cartoonish, monologue; average recording, no background noise; genuineness 2.6/6; vocal-burst blend 6.1/10; 16.9s, ITALIAN.
6807_7050_001476 · in -24.9 dBFS · gain +5.0 dB · mls-00078
(shame, embarrassment, pain · normal-paced, slightly relaxed, fairly steady, authoritative) e il poveretto si sentiva quel calcio nello stomaco bella provvidenza che avete gli diceva don franco il quale era venuto in maniche di camicia e col cappellaccio in testa a dare un'occhiata anche lui fumando la sua pipa
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, embarrassment, pain; style: authoritative, monologue; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 4.4/10; 13.3s, ITALIAN.
6807_7050_001200 · in -23.4 dBFS · gain +3.4 dB · mls-00078
Concentration ↓  /  Sournessidentity −0.05 emotion 96 %   c-mls-PXR · #19

This chain comes from the proxy rule: the same two-sided test as above, but because Sourness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Sourness clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.29.

At the same time Concentration goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.05, then +0.21, then +0.03 — a plateau around step 3, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 55 s · french · mls

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.945 before conversion and 0.896 after — it fell by 0.049. Neighbour-to-neighbour the worst pair went 0.923 → 0.846. (The earlier render, with segment 1 left raw, scores 0.845 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sourness moved +0.288 in the original and +0.276 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.263 became -0.291.

Quality. Mean predicted overall quality across the segments went 2.99 → 3.22 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.945 → 0.896 -0.049identity cos neighbours 0.923 → 0.846d_b rescored +0.288 → +0.276d_a rescored -0.263 → -0.291d_a mined -0.263d_b mined 0.288min_cos_consec (site) 0.9351min_cos_anchor (site) 0.9449dataset mlslang frenchspeaker 12709total 54.3schain gain +2.0 dBseam step 0.7 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, slightly relaxed, light breath
(concentration, contemplation · measured, normally alert, fairly steady, monologue) c'est que dans phybridité dans la variabilité notamment dans les variations simultanées appelées corrélation de croissance dans l'instinct dans les procédés de la concurrence vitale
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, contemplation; style: monologue, narration; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 1.6/10; 13.0s, FRENCH.
12709_14690_001040 · in -28.8 dBFS · gain +8.8 dB · mls-00059
(contemplation, concentration · measured, normally alert, steady, monologue) dans la sélection dans la succession géologique et dans la distribution géographique des êtres organisés dans les affinités mutuelles comme partout ailleurs la pensée de la nature est tatillonne et négligente économe et gâcheuse
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, concentration; style: monologue, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 15.2s, FRENCH.
12709_14690_000893 · in -27.2 dBFS · gain +7.2 dB · mls-00059
(awe, infatuation, emotional numbness · slow, subdued, steady, whispered) prévoyante et inattentive inconstante et inébranlable agitée et immobile une et innombrable grandiose et mesquine dans le môme moment et le même phénomène
full caption & clip details
An elderly somewhat feminine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, infatuation, emotional numbness; style: whispered, monologue; poor recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.6/10; 12.8s, FRENCH.
12709_14690_001333 · in -26.4 dBFS · gain +6.4 dB · mls-00060
(sourness, contempt, sadness · normal-paced, normally alert, steady, monologue) alors qu'elle avait devant elle le champs immense et vierge delà simplicité elle le peuple de petites erreurs de petites lois contradictoires de petits problèmes difficiles qui s'égarent dans rexislence comme des troupeaux aveugles
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, contempt, sadness; style: monologue, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.2/10; 14.0s, FRENCH.
12709_14690_000727 · in -27.6 dBFS · gain +7.6 dB · mls-00059
Sourness ↓  /  Concentrationidentity +0.00 emotion 68 %   c-mls-PXR · #20

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.29.

At the same time Sourness goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.12 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 49 s · dutch · mls

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.904 before conversion and 0.908 after — it rose by 0.003. Neighbour-to-neighbour the worst pair went 0.904 → 0.930. (The earlier render, with segment 1 left raw, scores 0.771 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.289 in the original and +0.198 after conversion — 68 % of the delta retained. On the other named axis, Sourness, -0.341 became -0.332.

Quality. Mean predicted overall quality across the segments went 3.22 → 3.35 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.904 → 0.908 +0.003identity cos neighbours 0.904 → 0.930d_b rescored +0.289 → +0.198d_a rescored -0.341 → -0.332d_a mined -0.341d_b mined 0.289min_cos_consec (site) 0.9246min_cos_anchor (site) 0.9246dataset mlslang dutchspeaker 2450total 48.1schain gain +1.1 dBseam step 1.2 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an elderly masculine voice · neutral-toned, slightly dark, balanced body, average recording, quiet background, frequent disfluency, slurred, audible breath
(sourness, disappointment, thankfulness gratitude · measured, very low-energy, slightly relaxed, monologue) op den derden dag konden zij al een beetje vliegen en nu dachten zij dat zij ook konden zweven en op de lucht drijven dat wilden zij maar bom daar duikelden zij daarom moesten zij hun vleugels gauw weer in beweging brengen
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, disappointment, thankfulness gratitude; style: monologue, narration; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.0/10; 19.2s, DUTCH.
2450_2211_000152 · in -27.1 dBFS · gain +7.1 dB · mls-00101
(measured, normally alert, slightly relaxed, storytelling) nu kwamen de jongens beneden op de straat en zongen hun lied ooievaar waar vlieg je heen zullen we niet naar beneden vliegen en hun de oogen uitpikken vroegen de jongen neen doet dat niet zei de moeder
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: storytelling, cartoonish; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 2.2/10; 15.8s, DUTCH.
2450_2211_000117 · in -27.8 dBFS · gain +7.8 dB · mls-00101
(concentration, contentment, triumph · slow, lethargic, relaxed, didactic) luistert maar naar mij dat is veel meer van belang een twee drie nu vliegen we rechts een twee drie nu links om den schoorsteen
full caption & clip details
An elderly masculine voice; delivery is lethargic, slow, relaxed, steady; timbre is neutral-toned, slightly dark, rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, contentment, triumph; style: didactic, whispered; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 0.2/10; 13.5s, DUTCH.
2450_2211_000078 · in -29.6 dBFS · gain +9.6 dB · mls-00101