This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_c-mls-B1.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words
are a real non-speech sound, printed where it happens.
The full generated caption for any clip is under “full caption & clip
details”. Its perceived-gender and background-noise clauses were re-rendered from
the numeric buckets, because the versions stored in the corpus index had those two
ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
The tags and captions describe the ORIGINAL clips — they are the
corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the
converted audio are the ones printed in each card's own paragraph and chips.
This chain comes from the one-sided rule: only Doubt had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Doubt clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.28.
Nothing was asked of the other axis, and in fact Hope Enthusiasm Optimism drifts down from 0.92 to 0.42 (-0.50), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.18, then +0.10 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 35 s · dutch · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.901 before conversion and 0.926 after — it rose by 0.025. Neighbour-to-neighbour the worst pair went 0.901 → 0.926. (The earlier render, with segment 1 left raw, scores 0.726 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.281 in the original and +0.380 after conversion — 135 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Hope Enthusiasm Optimism, -0.497 became -0.408.
Quality. Mean predicted overall quality across the segments went 3.24 → 3.38 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.901 → 0.926+0.025identity cos neighbours 0.901 → 0.926d_b rescored +0.281 → +0.380d_a rescored -0.497 → -0.408d_a mined -0.496d_b mined 0.281min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang dutchspeaker 1724total 33.9schain gain +2.5 dBseam step 1.1 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · fairly narrow pitch
(hope enthusiasm optimism, impatience and irritability · normal-paced, normally alert, slightly relaxed, whispered)ik zeg het eigenlijk enkel uit reactie op het zeurig en peuterig doen van berkie die zich verbeeldt naar documenten te werken wanneer ie enkele details heeft opgeteekend over menschen die hij kent daar verdoet ie dan dàgen aan
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; clear, almost no disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism, impatience and irritability; style: whispered, narration; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.7/10; 13.1s, DUTCH.
1724_2349_001627 · in -28.1 dBFS · gain +8.2 dB · mls-00102
(emotional numbness, confusion, fatigue exhaustion·measured, very low-energy, neutral tension, casual)hij heeft de tijd hè ja helaas jammer dat ie van de secretarie af is ik heb het er nog met hem over gehad z en is ie niet woedend geworden
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, slightly thin; average clarity, some disfluency, fairly narrow pitch, audible breath; affect is mildly positive, neutral stance, vulnerable; reads as emotional numbness, confusion, fatigue exhaustion; style: casual, conversational; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 3.7/10; 10.6s, DUTCH.
1724_2349_001610 · in -27.2 dBFS · gain +7.2 dB · mls-00102
(doubt, disappointment, confusion ·slow, very low-energy, slightly relaxed, whispered)och hij voelt het zelf ook wel hij verveelt zich is boos op zichzelf omdat ie niet voortkomt maar hij kan niet hij hééft niks te zeggen schrijf dàn eens raak
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, slightly relaxed, moderately variable; timbre is slightly cool, slightly dark, fairly smooth, thin; clear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, vulnerable; reads as doubt, disappointment, confusion; style: whispered, monologue; average recording, no background noise; genuineness 2.2/6; vocal-burst blend 0.4/10; 10.6s, DUTCH.
1724_2349_001616 · in -27.9 dBFS · gain +7.9 dB · mls-00102
This chain comes from the one-sided rule: only Affection had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Affection clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.38.
Nothing was asked of the other axis, and in fact Pain drifts down from 0.97 to 0.20 (-0.77), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.18 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 45 s · french · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.871 before conversion and 0.913 after — it rose by 0.042. Neighbour-to-neighbour the worst pair went 0.905 → 0.924. (The earlier render, with segment 1 left raw, scores 0.705 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Affection moved +0.379 in the original and +0.390 after conversion — 103 % of the delta retained, which is essentially all of it. On the other named axis, Pain, -0.773 became +0.473.
Quality. Mean predicted overall quality across the segments went 3.11 → 3.28 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.871 → 0.913+0.042identity cos neighbours 0.905 → 0.924d_b rescored +0.379 → +0.390d_a rescored -0.773 → +0.473d_a mined -0.773d_b mined 0.379min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang frenchspeaker 3698total 43.9schain gain +2.6 dBseam step 0.6 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult somewhat feminine voice · balanced body, average recording, quiet background, measured, slightly relaxed, fairly steady, fairly narrow pitch, light breath
(pain, awe, contemplation · subdued, no disfluency, clear, whispered)comme dans un miroir son caractère exhaussé sur le socle de la perversité idéale mais quand la fatigue impérieuse t'ordonnera d'arrêter ta marche devant les dalles de mon palais recouvertes de ronces et de chardons
full caption & clip details
A young adult somewhat feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as pain, awe, contemplation; style: whispered, monologue; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 2.0/10; 16.4s, FRENCH.
3698_1745_000069 · in -34.4 dBFS · gain +14.4 dB · mls-00062
(emotional numbness· subdued, frequent disfluency, slurred, whispered)fais attention à tes sandales en lambeaux et franchis sur la pointe des pieds l'élégance des vestibules ce n'est pas une recommandation inutile
full caption & clip details
A middle-aged somewhat feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as emotional numbness; style: whispered, monologue; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 2.7/10; 11.9s, FRENCH.
3698_1745_000083 · in -33.8 dBFS · gain +13.8 dB · mls-00062
(affection, teasing, malevolence malice·very low-energy, no disfluency, clear, monologue)tu pourrais éveiller ma jeune épouse et mon fils en bas âge couchés dans les caveaux de plomb qui longent les fondements de l'antique château si tu ne prenais tes précautions d'avance ils pourraient te faire pâlir par leurs hurlements souterrains
full caption & clip details
A young adult somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, neutral openness; reads as affection, teasing, malevolence malice; style: monologue, narration; average recording, quiet background; explicit content; genuineness 1.4/6; vocal-burst blend 2.4/10; 16.0s, FRENCH.
3698_1745_000008 · in -35.6 dBFS · gain +15.6 dB · mls-00062
This chain comes from the one-sided rule: only Malevolence Malice had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Malevolence Malice clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.30.
Nothing was asked of the other axis, and in fact Concentration drifts down from 0.96 to 0.80 (-0.16), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.11 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 41 s · german · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.839 before conversion and 0.895 after — it rose by 0.056. Neighbour-to-neighbour the worst pair went 0.882 → 0.919. (The earlier render, with segment 1 left raw, scores 0.809 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Malevolence Malice moved +0.296 in the original and +0.286 after conversion — 97 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.162 became -0.097.
Quality. Mean predicted overall quality across the segments went 3.09 → 3.25 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.839 → 0.895+0.056identity cos neighbours 0.882 → 0.919d_b rescored +0.296 → +0.286d_a rescored -0.162 → -0.097d_a mined -0.162d_b mined 0.296min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang germanspeaker 10148total 40.0schain gain +3.2 dBseam step 1.5 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(concentration · normal-paced, steady, no disfluency, monologue)die unterhaltung sprang nun sofort auf das neue thema der frauenbildung über alexei alexandrowitsch sprach den gedanken aus daß die frage der frauenbildung gewöhnlich mit der frage der frauengleichberechtigung zusammengeworfen werde
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.3/10; 15.6s, GERMAN.
10148_10349_007213 · in -28.5 dBFS · gain +8.5 dB · mls-00002
(sourness, emotional numbness, fear·measured, steady, no disfluency, didactic)und daß man die frauenbildung nur deshalb für schädlich erachte ich bin im gegenteil der ansicht daß diese beiden fragen unlösbar miteinander verknüpft sind sagte peszow
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, emotional numbness, fear; style: didactic, narration; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.3/10; 13.2s, GERMAN.
10148_10349_007052 · in -27.1 dBFS · gain +7.1 dB · mls-00002
(malevolence malice, contempt, fear · measured, fairly steady, almost no disfluency, didactic)es liegt da ein irreführender zirkelschluß vor der frau werden die rechte versagt wegen des mangels an bildung und der mangel an bildung ist eine folge des fehlens von rechten
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, contempt, fear; style: didactic, formal; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.2/10; 11.6s, GERMAN.
10148_10349_006448 · in -26.0 dBFS · gain +6.0 dB · mls-00001
This chain comes from the one-sided rule: only Relief had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Relief clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.27.
Nothing was asked of the other axis, and in fact Malevolence Malice barely moves at all, sitting near 0.89 throughout.
It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.02, then +0.01 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 59 s · french · mls
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.940 before conversion and 0.916 after — it fell by 0.024. Neighbour-to-neighbour the worst pair went 0.943 → 0.916. (The earlier render, with segment 1 left raw, scores 0.802 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.267 in the original and +0.210 after conversion — 79 % of the delta retained, which is most of it. On the other named axis, Malevolence Malice, -0.001 became -0.019.
Quality. Mean predicted overall quality across the segments went 3.23 → 3.29 (+0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.940 → 0.916-0.024identity cos neighbours 0.943 → 0.916d_b rescored +0.267 → +0.210d_a rescored -0.001 → -0.019d_a mined -0.001d_b mined 0.267min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang frenchspeaker 3698total 57.5schain gain +1.6 dBseam step 0.6 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult somewhat feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, slightly relaxed, light breath
(measured, subdued, steady, monologue)le grand prêtre qui a le pouvoir d'être écouté des plus grands juges a dit et fait des choses que nous ne savons point puis il a mandé mon frère devant lui et lui a dit mon fils confessez-vous à moi comme à dieu
full caption & clip details
A young adult somewhat feminine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: monologue, whispered; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.1/10; 14.1s, FRENCH.
3698_1902_000384 · in -30.9 dBFS · gain +10.9 dB · mls-00063
(distress, fear, malevolence malice· measured, subdued, steady, monologue)et huriel ayant dit toute la vérité de bout en bout l'évêque lui a dit encore faites-en pénitence mon fils et repentez-vous votre affaire est arrangée devant les hommes vous n'en serez jamais inquiété
full caption & clip details
A young adult somewhat feminine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as distress, fear, malevolence malice; style: monologue, whispered; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.9/10; 12.8s, FRENCH.
3698_1902_000605 · in -30.1 dBFS · gain +10.2 dB · mls-00063
(contentment, thankfulness gratitude, relief·normal-paced, normally alert, steady, monologue)mais vous devez apaiser le mécontentement de dieu et pour cela je vous engage à quitter la compagnie et la confrérie des muletiers qui sont gens sans religion et dont les pratiques secrètes sont contraires aux lois du ciel et de la terre
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contentment, thankfulness gratitude, relief; style: monologue, storytelling; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.6/10; 13.3s, FRENCH.
3698_1902_000342 · in -31.2 dBFS · gain +11.2 dB · mls-00063
(relief, sadness, disappointment·measured, subdued, fairly steady, monologue)et mon frère lui ayant humblement remontré qu'il s'y trouvait pourtant d'honnêtes gens c'est tant pis a dit le grand prêtre si les honnêtes gens qui s'y trouvent refusaient les serments qui s'y font le mal sortirait de cette société là et ce serait une corporation d'ouvriers aussi estimable que toute autre
full caption & clip details
A young adult somewhat feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as relief, sadness, disappointment; style: monologue, storytelling; good recording, quiet background; genuineness 1.0/6; vocal-burst blend 2.4/10; 17.9s, FRENCH.
3698_1902_000357 · in -31.0 dBFS · gain +11.0 dB · mls-00063
This chain comes from the one-sided rule: only Contemplation had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Contemplation clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.22.
Nothing was asked of the other axis, and in fact Sexual Lust drifts down from 0.99 to 0.28 (-0.71), which the rule did not require.
It takes 2 clips to get there. Clip to clip the moves are +0.22 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 27 s · german · mls
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.862 before conversion and 0.903 after — it rose by 0.041. Neighbour-to-neighbour the worst pair went 0.862 → 0.903. (The earlier render, with segment 1 left raw, scores 0.850 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.224 in the original and +0.285 after conversion — 127 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Sexual Lust, -0.707 became -0.983.
Quality. Mean predicted overall quality across the segments went 3.11 → 3.20 (+0.08) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.862 → 0.903+0.041identity cos neighbours 0.862 → 0.903d_b rescored +0.224 → +0.285d_a rescored -0.707 → -0.983d_a mined -0.707d_b mined 0.224min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang germanspeaker 9565total 26.8schain gain +2.2 dBseam step 0.4 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, no background noise, measured, fairly steady, light breath
(sexual lust, longing, affection · highly aroused, neutral tension, almost no disfluency, storytelling)gleichwohl hoffe ich daß du in deiner günstigen stimmung gegen mich bestärkt werden wirst wenn ich um deinem befehle zu genügen dir mein abenteuer erzählt haben werde
full caption & clip details
A middle-aged masculine voice; delivery is highly aroused, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, rough, thin; very clear, almost no disfluency, wide pitch range, light breath; affect is neutral, neutral stance, fairly guarded; reads as sexual lust, longing, affection; style: storytelling, narration; average recording, no background noise; genuineness 0.5/6; vocal-burst blend 2.6/10; 13.1s, GERMAN.
9565_10506_002709 · in -26.7 dBFS · gain +6.7 dB · mls-00004
(contemplation, fear, longing ·normally alert, slightly relaxed, little disfluency, monologue)nach dieser höflichen anrede wodurch er sich des wohlwollens und der aufmerksamkeit des kalifen versichern wollte und nachdem er sich noch einige augenblicke das was er zu sagen hatte in sein gedächtnis zurückgerufen
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, fear, longing; style: monologue, narration; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.9/10; 14.0s, GERMAN.
9565_10506_001932 · in -26.9 dBFS · gain +6.9 dB · mls-00004
This chain comes from the one-sided rule: only Fear had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Fear clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.25.
Nothing was asked of the other axis, and in fact Longing barely moves at all, sitting near 0.99 throughout.
It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.09 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 50 s · german · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.933 before conversion and 0.920 after — it fell by 0.013. Neighbour-to-neighbour the worst pair went 0.956 → 0.927. (The earlier render, with segment 1 left raw, scores 0.798 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.253 in the original and +0.483 after conversion — 191 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Longing, -0.048 became -0.028.
Quality. Mean predicted overall quality across the segments went 3.12 → 3.22 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.933 → 0.920-0.013identity cos neighbours 0.956 → 0.927d_b rescored +0.253 → +0.483d_a rescored -0.048 → -0.028d_a mined -0.048d_b mined 0.253min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang germanspeaker 9565total 49.4schain gain +1.4 dBseam step 0.8 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, balanced body, slightly relaxed, clear, light breath
(longing, sourness, malevolence malice · normal-paced, normally alert, fairly steady, newsreading)wisse auch mahmud fügte noch mein lehrer hinzu deine brüder haben an der tür alles gehört was ich dir bisher gesagt und lassen sich in diesem augenblick von zwei geistern nach ägypten bringen denn sie glauben wenn sie das befolgen was ich dir anempfohlen habe statt deiner sich das schwert und das buch zueignen zu können
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing, sourness, malevolence malice; style: newsreading, narration; average recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.0/10; 20.0s, GERMAN.
9565_13024_003370 · in -25.1 dBFS · gain +5.1 dB · mls-00013
(awe, bitterness, helplessness·measured, subdued, steady, narration)aber sowie sie in den see karun steigen werden sie von den genien des sees getötet doch nur gott ist allwissend nach diesen worten rief mein lehrer den geist der mich von dem kloster nach tunis gebracht hatte und befahl ihm mich nach ägypten zu tragen
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; clear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, bitterness, helplessness; style: narration, monologue; average recording, quiet background; genuineness 0.8/6; vocal-burst blend 0.0/10; 17.3s, GERMAN.
9565_13024_002727 · in -26.9 dBFS · gain +6.9 dB · mls-00013
(fear, awe, sexual lust· measured, normally alert, fairly steady, narration)der geist breitete sogleich seine flügel aus und trug mich bis in die nähe des sees karun dann verschwand er und brachte mir eine djinn in der gestalt eines maultiers und setzte mich darauf
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, awe, sexual lust; style: narration, storytelling; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.3/10; 12.6s, GERMAN.
9565_13024_003435 · in -28.1 dBFS · gain +8.1 dB · mls-00013
This chain comes from the one-sided rule: only Doubt had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Doubt clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.33.
Nothing was asked of the other axis, and in fact Malevolence Malice barely moves at all, sitting near 0.93 throughout.
It takes 4 clips to get there. Clip to clip the moves are +0.21, then +0.00, then +0.12 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 66 s · dutch · mls
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.930 before conversion and 0.883 after — it fell by 0.047. Neighbour-to-neighbour the worst pair went 0.930 → 0.883. (The earlier render, with segment 1 left raw, scores 0.735 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.333 in the original and +0.151 after conversion — 45 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Malevolence Malice, -0.029 became -0.084.
Quality. Mean predicted overall quality across the segments went 3.08 → 3.53 (+0.45) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.930 → 0.883-0.047identity cos neighbours 0.930 → 0.883d_b rescored +0.333 → +0.151d_a rescored -0.029 → -0.084d_a mined -0.029d_b mined 0.333min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang dutchspeaker 2450total 64.8schain gain +1.0 dBseam step 1.1 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-toned, slightly rough, balanced body, average recording, quiet background, measured, slightly relaxed, fairly steady
(malevolence malice, sexual lust · normally alert, some disfluency, somewhat unclear, cartoonish)dien brief had ik niet moeten hebben zeide de witt half aarzelende of hy hem al dan niet zoû openen reden te meer om hem te lezen merkte van espenblad aan
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as malevolence malice, sexual lust; style: cartoonish, narration; average recording, quiet background; genuineness 0.6/6; vocal-burst blend 0.6/10; 12.2s, DUTCH.
2450_10026_000409 · in -29.5 dBFS · gain +9.5 dB · mls-00084
(triumph, concentration, anger·subdued, frequent disfluency, slurred, cartoonish)wat niet voor uw oog bestemd was zal zeker belangrijker zijn dan wat men u wel wil mededeelen ik beken u gul weg hernam de witt dat het denkbeeld my stuit om op deze wijze misbruik van iemands goed vertrouwen te maken en in zijn geheimen te dringen
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, concentration, anger; style: cartoonish, monologue; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.4/10; 19.1s, DUTCH.
2450_10026_000796 · in -28.2 dBFS · gain +8.2 dB · mls-00084
(triumph, concentration, anger · subdued, frequent disfluency, slurred, cartoonish)wat niet voor uw oog bestemd was zal zeker belangrijker zijn dan wat men u wel wil mededeelen ik beken u gul weg hernam de witt dat het denkbeeld my stuit om op deze wijze misbruik van iemands goed vertrouwen te maken en in zijn geheimen te dringen
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, concentration, anger; style: cartoonish, monologue; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.4/10; 19.1s, DUTCH.
2450_10026_000796 · in -28.2 dBFS · gain +8.2 dB · mls-00084
(doubt, contempt, concentration ·normally alert, some disfluency, average clarity, cartoonish)het hart van van espenblad begon van angst te popelen by de gedachte dat een brief waaruit naar zijn stellige overtuiging een aanklacht tegen buat te halen ware door de witt ongelezen zoû worden terug
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, contempt, concentration; style: cartoonish, monologue; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 0.9/10; 15.0s, DUTCH.
2450_10026_003761 · in -28.4 dBFS · gain +8.4 dB · mls-00084
This chain comes from the one-sided rule: only Shame had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Shame clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.29.
Nothing was asked of the other axis, and in fact Sourness drifts down from 0.97 to 0.81 (-0.16), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.04 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 47 s · german · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.875 before conversion and 0.904 after — it rose by 0.029. Neighbour-to-neighbour the worst pair went 0.925 → 0.932. (The earlier render, with segment 1 left raw, scores 0.783 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.287 in the original and +0.059 after conversion — 21 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Sourness, -0.162 became -0.130.
Quality. Mean predicted overall quality across the segments went 3.07 → 3.28 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.875 → 0.904+0.029identity cos neighbours 0.925 → 0.932d_b rescored +0.287 → +0.059d_a rescored -0.162 → -0.130d_a mined -0.162d_b mined 0.286min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang germanspeaker 2037total 46.1schain gain +1.9 dBseam step 0.8 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, measured, normally alert, slightly relaxed, fairly steady
(sourness, disappointment, anger · little disfluency, monologue, didactic)er versprach ihr für ihr weiteres fortkommen in ihre heimat zu sorgen wenn sie ihm behülflich sein wollte ins schloß zu gelangen aber ein gedanke machte ihm noch sorge nämlich der woher er zwei oder drei treue gehülfen bekommen könnte
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, disappointment, anger; style: monologue, didactic; average recording, quiet background; genuineness 0.5/6; vocal-burst blend 0.0/10; 18.3s, GERMAN.
2037_2034_000820 · in -21.9 dBFS · gain +1.9 dB · mls-00019
(bitterness, thankfulness gratitude, relief·no disfluency, monologue, formal)da fiel ihm orbasans dolch ein und das versprechen das ihm jener gegeben hatte ihm wo er seiner bedürfe zu hülfe zu eilen und er machte sich daher mit fatme aus dem begräbnis auf um den räuber aufzusuchen
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as bitterness, thankfulness gratitude, relief; style: monologue, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 16.4s, GERMAN.
2037_2034_001130 · in -22.2 dBFS · gain +2.2 dB · mls-00019
(shame, emotional numbness, sadness· no disfluency, formal, monologue)in der nämlichen stadt wo er sich zum arzt umgewandelt hatte kaufte er um sein letztes geld ein roß und mietete fatme bei einer armen frau in der vorstadt ein
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, emotional numbness, sadness; style: formal, monologue; very good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.8s, GERMAN.
2037_2034_000105 · in -22.8 dBFS · gain +2.8 dB · mls-00018
This chain comes from the one-sided rule: only Contemplation had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Contemplation strongly present — 0.78, higher than 78 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.21.
Nothing was asked of the other axis, and in fact Sadness drifts down from 0.98 to 0.43 (-0.55), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.00, then +0.21 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 57 s · french · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.940 before conversion and 0.924 after — it fell by 0.016. Neighbour-to-neighbour the worst pair went 0.940 → 0.928. (The earlier render, with segment 1 left raw, scores 0.786 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.210 in the original and +0.256 after conversion — 122 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Sadness, -0.547 became -0.516.
Quality. Mean predicted overall quality across the segments went 3.21 → 3.32 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.940 → 0.924-0.016identity cos neighbours 0.940 → 0.928d_b rescored +0.210 → +0.256d_a rescored -0.547 → -0.516d_a mined -0.547d_b mined 0.210min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang frenchspeaker 1840total 56.4schain gain -0.7 dBseam step 0.9 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · slightly warm, slightly dark, slightly rough, balanced body, average recording, quiet background, measured, steady
(thankfulness gratitude, sadness, disappointment · subdued, slightly relaxed, somewhat unclear, monologue)maternelle alice lui dit-elle le huron nous offre la vie à toutes deux il fait plus il promet de vous rendre vous et notre cher duncan à la liberté à nos amis à notre malheureux père si je puis dompter ce cœur rebelle cet orgueil de fierté au point de consentir
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, sadness, disappointment; style: monologue, storytelling; average recording, quiet background; genuineness 0.6/6; vocal-burst blend 2.2/10; 19.2s, FRENCH.
1840_10379_003828 · in -30.6 dBFS · gain +10.6 dB · mls-00044
(thankfulness gratitude, sadness, disappointment · subdued, slightly relaxed, somewhat unclear, monologue)maternelle alice lui dit-elle le huron nous offre la vie à toutes deux il fait plus il promet de vous rendre vous et notre cher duncan à la liberté à nos amis à notre malheureux père si je puis dompter ce cœur rebelle cet orgueil de fierté au point de consentir
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, sadness, disappointment; style: monologue, storytelling; average recording, quiet background; genuineness 0.6/6; vocal-burst blend 2.2/10; 19.2s, FRENCH.
1840_10379_003828 · in -30.6 dBFS · gain +10.6 dB · mls-00044
(contemplation, awe, longing·very low-energy, relaxed, slurred, whispered)consentir la voix lui manqua elle joignit les mains et leva les yeux vers le ciel comme pour supplier la sagesse infinie de lui inspirer ce qu'elle devait dire ce qu'elle devait faire de consentir à quoi s'écria alice continuez ma chère cora quexige t il de nous
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, relaxed, steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, awe, longing; style: whispered, monologue; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 2.7/10; 18.4s, FRENCH.
1840_10379_003757 · in -31.7 dBFS · gain +11.7 dB · mls-00044
This chain comes from the one-sided rule: only Infatuation had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Infatuation clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.21.
Nothing was asked of the other axis, and in fact Confusion drifts down from 0.97 to 0.87 (-0.10), which the rule did not require.
It takes 2 clips to get there. Clip to clip the moves are +0.21 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 24 s · dutch · mls
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.915 before conversion and 0.875 after — it fell by 0.041. Neighbour-to-neighbour the worst pair went 0.915 → 0.875. (The earlier render, with segment 1 left raw, scores 0.763 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
The emotional move did not survive. Re-scored end to end, Infatuation moved +0.211 in the original and -0.055 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Confusion, -0.097 became -0.051.
Quality. Mean predicted overall quality across the segments went 3.15 → 3.37 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.915 → 0.875-0.041identity cos neighbours 0.915 → 0.875d_b rescored +0.211 → -0.055d_a rescored -0.097 → -0.051d_a mined -0.097d_b mined 0.211min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang dutchspeaker 2450total 23.8schain gain +1.5 dBseam step 0.1 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an elderly masculine voice · neutral-toned, slightly dark, slightly rough, average recording, no background noise, slightly relaxed, fairly narrow pitch
(confusion, fatigue exhaustion, sourness · slow, very low-energy, steady, narration)trouwelooze en gij maar neen mauritio niet trouwloos zijt gij geweest ik ongelukkige ik heb uwe boert voor ernst gehouden
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, fatigue exhaustion, sourness; style: narration, monologue; average recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.2/10; 11.2s, DUTCH.
2450_10565_001443 · in -28.4 dBFS · gain +8.4 dB · mls-00088
(infatuation, concentration, sexual lust·measured, normally alert, fairly steady, narration)maar met een verschrikkelijken lach op haar aangezigt maar gij boert gij scherst weder er is immers geen zulke maria zeide gij niet dat gij nooit een meisje gezien hadt dat mij overtrof
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, full; clear, almost no disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation, concentration, sexual lust; style: narration, storytelling; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.3/10; 12.9s, DUTCH.
2450_10565_001247 · in -29.1 dBFS · gain +9.1 dB · mls-00088
This chain comes from the one-sided rule: only Thankfulness Gratitude had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Thankfulness Gratitude strongly present — 0.78, higher than 78 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.22.
Nothing was asked of the other axis, and in fact Fatigue Exhaustion drifts down from 1.00 to 0.79 (-0.21), which the rule did not require.
It takes 2 clips to get there. Clip to clip the moves are +0.22 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 25 s · german · mls
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.952 before conversion and 0.924 after — it fell by 0.028. Neighbour-to-neighbour the worst pair went 0.952 → 0.924. (The earlier render, with segment 1 left raw, scores 0.867 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.218 in the original and +0.259 after conversion — 119 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Fatigue Exhaustion, -0.212 became -0.091.
Quality. Mean predicted overall quality across the segments went 2.98 → 3.17 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.952 → 0.924-0.028identity cos neighbours 0.952 → 0.924d_b rescored +0.218 → +0.259d_a rescored -0.212 → -0.091d_a mined -0.212d_b mined 0.218min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang germanspeaker 10148total 24.4schain gain +3.3 dBseam step 0.6 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an elderly somewhat feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, slightly relaxed, little disfluency
(fatigue exhaustion, longing, fear · measured, subdued, fairly steady, narration)kaum zwölfjährig half ich schon in den stunden die mir die schule freiließ im haushalt in der küche bei der wäsche ich tat auch alles recht gern es fiel mir gar nicht ein daß es anders hätte sein können
full caption & clip details
An elderly somewhat feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, neutral openness; reads as fatigue exhaustion, longing, fear; style: narration, whispered; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.9/10; 14.1s, GERMAN.
10148_11805_000516 · in -28.1 dBFS · gain +8.1 dB · mls-00010
(thankfulness gratitude, affection, longing ·slow, very low-energy, variable, whispered)alle mädchen die wie wir in einfachen verhältnissen lebten taten so ziemlich dasselbe ich war heiter zufrieden und kerngesund
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, slow, slightly relaxed, variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, fairly narrow pitch, audible breath; affect is mildly positive, neutral stance, vulnerable; reads as thankfulness gratitude, affection, longing; style: whispered, ASMR; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 1.2/10; 10.4s, GERMAN.
10148_11805_000135 · in -29.1 dBFS · gain +9.2 dB · mls-00010
This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Emotional Numbness clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.26.
Nothing was asked of the other axis, and in fact Bitterness drifts down from 1.00 to 0.74 (-0.26), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.11, then +0.15 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 54 s · spanish · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.936 before conversion and 0.949 after — it rose by 0.014. Neighbour-to-neighbour the worst pair went 0.939 → 0.938. (The earlier render, with segment 1 left raw, scores 0.785 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.263 in the original and +0.324 after conversion — 123 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Bitterness, -0.271 became -0.102.
Quality. Mean predicted overall quality across the segments went 2.97 → 3.22 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.936 → 0.949+0.014identity cos neighbours 0.939 → 0.938d_b rescored +0.263 → +0.324d_a rescored -0.271 → -0.102d_a mined -0.261d_b mined 0.263min_cos_consec (site) 0.9388min_cos_anchor (site) 0.9388dataset mlslang spanishspeaker 10246total 53.7schain gain +2.5 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, almost no disfluency, clear
(bitterness, anger, sourness · measured, cartoonish, narration)así aun cuando tuviese al mismo tiempo vivos deseos de arrojarle por la ventana reprimióse muy cuerdamente y dijo rosalía de castro qué puedo sacar en consecuencia de cuanto acaba usted de decir habló usted únicamente para sí mismo ó para los dos para mí mismo y para los dos
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as bitterness, anger, sourness; style: cartoonish, narration; average recording, no background noise; genuineness 1.7/6; vocal-burst blend 4.2/10; 19.9s, SPANISH.
10246_11789_000691 · in -26.3 dBFS · gain +6.3 dB · mls-00037
(sourness, anger, shame·normal-paced, monologue, narration)pues expliqúese usted más claro y pronto en todo lo que haya dicho para mí porque no he comprendido nada y necesito comprender quise decir para los dos que nuestras singularidades se atraen como dos cuerpos simpáticos y si no dígalo la breve cuanto larga historia de nuestra entrevista
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, anger, shame; style: monologue, narration; good recording, quiet background; genuineness 0.8/6; vocal-burst blend 2.3/10; 19.6s, SPANISH.
10246_11789_000341 · in -26.6 dBFS · gain +6.6 dB · mls-00037
(emotional numbness, shame, disgust· normal-paced, narration, monologue)llego me anuncio sin preámbulos como persona que sabe no han de negarle el paso usted sale á recibirme nada de eso á arrojarle á usted aun cuando fuese por la ventana si se hiciese preciso
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, shame, disgust; style: narration, monologue; good recording, quiet background; genuineness 1.0/6; vocal-burst blend 1.7/10; 14.6s, SPANISH.
10246_11789_000664 · in -26.7 dBFS · gain +6.7 dB · mls-00037
This chain comes from the one-sided rule: only Fear had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Fear clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.26.
Nothing was asked of the other axis, and in fact Emotional Numbness barely moves at all, sitting near 0.96 throughout.
It takes 3 clips to get there. Clip to clip the moves are +0.08, then +0.18 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 44 s · spanish · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.900 before conversion and 0.886 after — it fell by 0.014. Neighbour-to-neighbour the worst pair went 0.932 → 0.947. (The earlier render, with segment 1 left raw, scores 0.783 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.261 in the original and +0.595 after conversion — 228 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, +0.013 became -0.014.
Quality. Mean predicted overall quality across the segments went 2.79 → 3.18 (+0.40) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.900 → 0.886-0.014identity cos neighbours 0.932 → 0.947d_b rescored +0.261 → +0.595d_a rescored +0.013 → -0.014d_a mined 0.013d_b mined 0.261min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang spanishspeaker 10246total 43.4schain gain +0.5 dBseam step 3.6 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(emotional numbness, jealousy and envy, pride · measured, almost no disfluency, narration, monologue)en todo caso resolvió modificar su testamento dejando la herencia de cati no en sus manos sino en las de otros herederos personas de confianza concediéndole sólo el usufructo y luego la plena posesión a sus hijos caso de que los tuviera
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, jealousy and envy, pride; style: narration, monologue; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 2.5/10; 18.0s, SPANISH.
10246_10857_002544 · in -27.1 dBFS · gain +7.2 dB · mls-00029
(malevolence malice·normal-paced, no disfluency, narration, monologue)así los bienes de catalina no irían a heathcliff aunque muriese su hijo de acuerdo con sus instrucciones envié a un hombre en busca del procurador y a otros cuatro estos armados a buscar a la señorita
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice; style: narration, monologue; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.2/10; 14.3s, SPANISH.
10246_10857_002397 · in -27.0 dBFS · gain +7.0 dB · mls-00029
(fear, disappointment, emotional numbness· normal-paced, no disfluency, monologue, authoritative)el primero de ellos volvió anunciando que había tenido que estar dos horas esperando al señor green y que éste vendría al siguiente día ya que tenía que hacer en el pueblo
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, disappointment, emotional numbness; style: monologue, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.0/10; 11.4s, SPANISH.
10246_10857_002171 · in -25.5 dBFS · gain +5.5 dB · mls-00029
This chain comes from the one-sided rule: only Longing had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Longing strongly present — 0.78, higher than 78 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.21.
Nothing was asked of the other axis, and in fact Shame barely moves at all, sitting near 0.95 throughout.
It takes 3 clips to get there. Clip to clip the moves are +0.09, then +0.12 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 41 s · dutch · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.884 before conversion and 0.905 after — it rose by 0.021. Neighbour-to-neighbour the worst pair went 0.884 → 0.905. (The earlier render, with segment 1 left raw, scores 0.740 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.207 in the original and +0.147 after conversion — 71 % of the delta retained, which is most of it. On the other named axis, Shame, -0.028 became -0.107.
Quality. Mean predicted overall quality across the segments went 3.26 → 3.46 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.884 → 0.905+0.021identity cos neighbours 0.884 → 0.905d_b rescored +0.207 → +0.147d_a rescored -0.028 → -0.107d_a mined -0.028d_b mined 0.207min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang dutchspeaker 1724total 40.6schain gain +1.9 dBseam step 0.7 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · neutral-toned, neutral-bright, fairly smooth, normally alert, slightly relaxed, fairly steady, light breath
(shame · measured, some disfluency, average clarity, monologue)dit l'autre avec les loups en ik mag van baalen wel waarschuwen anders gaat gij nog met de kas strijken en dan zou hij inderdaad de ongelukkigste man der wereld worden
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as shame; style: monologue, narration; good recording, quiet background; genuineness 0.5/6; vocal-burst blend 0.0/10; 15.3s, DUTCH.
1724_1601_002362 · in -28.2 dBFS · gain +8.2 dB · mls-00096
(astonishment surprise·normal-paced, no disfluency, clear, narration)ik heb al aan jetje blaek geraden haar koffertje na te kijken om te zien of gij haar niets ontstolen hebt ten haren opzichte zou ik het u vergeven want dat ware volgens het recht van
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as astonishment surprise; style: narration, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 10.8s, DUTCH.
1724_1601_001979 · in -29.1 dBFS · gain +9.1 dB · mls-00096
(longing, contemplation, disappointment· normal-paced, no disfluency, clear, whispered)zou papa zeggen ja ik weet ook wel een mondvol latijn overmits zij zoo t mij voorkomt zich aan de dieverij van uw hart heeft schuldig gemaakt nu op haar hart zult gij wel alle aanspraak verbeuren zoo gij er dergelijke kennissen op nahoudt
full caption & clip details
A young adult somewhat feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; clear, no disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as longing, contemplation, disappointment; style: whispered, narration; average recording, quiet background; genuineness 0.9/6; vocal-burst blend 1.6/10; 14.9s, DUTCH.
1724_1601_001608 · in -28.9 dBFS · gain +8.9 dB · mls-00096
This chain comes from the one-sided rule: only Pride had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Pride clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.26.
Nothing was asked of the other axis, and in fact Malevolence Malice barely moves at all, sitting near 0.93 throughout.
It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.01, then +0.09, then -0.01 — not a clean run: step 4 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 78 s · spanish · mls
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.935 before conversion and 0.877 after — it fell by 0.058. Neighbour-to-neighbour the worst pair went 0.960 → 0.924. (The earlier render, with segment 1 left raw, scores 0.808 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.258 in the original and +0.219 after conversion — 85 % of the delta retained, which is most of it. On the other named axis, Malevolence Malice, -0.002 became +0.013.
Quality. Mean predicted overall quality across the segments went 3.18 → 3.34 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.935 → 0.877-0.058identity cos neighbours 0.960 → 0.924d_b rescored +0.258 → +0.219d_a rescored -0.002 → +0.013d_a mined -0.002d_b mined 0.258min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang spanishspeaker 3946total 76.9schain gain -1.5 dBseam step 0.5 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normal-paced, normally alert, slightly relaxed
(malevolence malice · fairly steady, no disfluency, moderate pitch range, monologue)dexemos de hablar des to del duero y diré como cortés luego mandó llamar á un nuestro capitan que se dice juan velazquez de leon persona de mucha cuenta y amigo de cortés y era pariente muy cercano del gobernador de cuba diego velazquez
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice; style: monologue, narration; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 2.4/10; 15.2s, SPANISH.
3946_11219_004886 · in -25.0 dBFS · gain +5.0 dB · mls-00032
(thankfulness gratitude, sourness, sadness· fairly steady, almost no disfluency, moderate pitch range, monologue)á lo que siempre tuvimos creido tambien le tenia cortés convocado y atraido á si con grandes dádivas y ofrecimientos que le darla mando en la nueva españa y le haria su igual porque el juan velazquez siempre se mostró muy gran servidor y verdadero amigo como en adelante verán
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, sourness, sadness; style: monologue, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 4.5/10; 18.8s, SPANISH.
3946_11219_005345 · in -25.7 dBFS · gain +5.7 dB · mls-00032
(sourness, malevolence malice, fear·steady, no disfluency, moderate pitch range, monologue)a lo que señor juan velazquez le hice llamar es que me dixo andres de duero que dice narvaez y en todo su real hay fama que si v m va allá que luego yo soy deshecho y desbaratado porque creen que se ha de hacer con narvaez
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, malevolence malice, fear; style: monologue, narration; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 2.7/10; 16.2s, SPANISH.
3946_11219_005195 · in -25.9 dBFS · gain +5.8 dB · mls-00032
(sourness, bitterness, shame· steady, no disfluency, moderate pitch range, monologue)y á esta causa he acordado que por mi vida si bien me quiere que luego se vaya en su buena yegua rucia y que lleve todo su oro y la fanfarrona que era muy pesada cadena de oro y otras cositas que yo le daré que dé allá por mi á quien yo le dixere
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, bitterness, shame; style: monologue, narration; good recording, quiet background; genuineness 0.1/6; vocal-burst blend 2.6/10; 16.8s, SPANISH.
3946_11219_004244 · in -25.1 dBFS · gain +5.1 dB · mls-00032
(pride, sadness, longing· steady, no disfluency, fairly narrow pitch, monologue)dixere su fanfarrona de oro que pesaba mucho llevará al hombro y otra cadena que pesa mas que ella llevara con dos vueltas y allá verá que le quiere narvaez
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, sadness, longing; style: monologue, narration; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 10.6s, SPANISH.
3946_11219_004275 · in -26.4 dBFS · gain +6.4 dB · mls-00032
This chain comes from the one-sided rule: only Contempt had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Contempt strongly present — 0.77, higher than 77 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.21.
Nothing was asked of the other axis, and in fact Contemplation barely moves at all, sitting near 0.97 throughout.
It takes 3 clips to get there. Clip to clip the moves are -0.00, then +0.21 — not a clean run: step 1 moves back the other way by 0.00 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 50 s · french · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.931 before conversion and 0.894 after — it fell by 0.038. Neighbour-to-neighbour the worst pair went 0.936 → 0.900. (The earlier render, with segment 1 left raw, scores 0.756 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contempt moved +0.207 in the original and +0.157 after conversion — 76 % of the delta retained, which is most of it. On the other named axis, Contemplation, -0.029 became +0.041.
Quality. Mean predicted overall quality across the segments went 3.04 → 3.16 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.931 → 0.894-0.038identity cos neighbours 0.936 → 0.900d_b rescored +0.207 → +0.157d_a rescored -0.029 → +0.041d_a mined -0.029d_b mined 0.207min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang frenchspeaker 12709total 49.3schain gain +1.1 dBseam step 1.9 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, normally alert, slightly relaxed, steady, no disfluency, clear
(contemplation, jealousy and envy, concentration · normal-paced, formal, monologue)il est amateur de courses et volontiers spectateur de départs d'aviation il est sauf quand il est atteint de paresse physique très grand voyageur les voyages étant sinon tout à fait comme a dit emerson le paradis des sots du moins le paradis de tous ceux à qui le don d'observer ou de méditer est refusé
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, jealousy and envy, concentration; style: formal, monologue; good recording, quiet background; genuineness 0.1/6; vocal-burst blend 0.6/10; 19.0s, FRENCH.
12709_13651_000897 · in -26.0 dBFS · gain +6.0 dB · mls-00054
(fatigue exhaustion, pain, longing·measured, whispered, narration)ni la méditation ni même l'observation ne demandant plus de six kilomètres carrés pour se satisfaire il est très volontiers conteur et conteur de soi-même il est celui qui dit le plus j'étais là telle chose
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, pain, longing; style: whispered, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.1/10; 15.0s, FRENCH.
12709_13651_000210 · in -25.3 dBFS · gain +5.3 dB · mls-00054
(contempt, contemplation, pride· measured, monologue, didactic)il conte beaucoup raisonne peu ne réfléchit jamais et ignore le repentir c'est un homme aimable dont la société est aussi agréable qu'elle est inutile s'il est vrai ce que l'on pourra contester que ce qui est agréable puisse être inutile
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, contemplation, pride; style: monologue, didactic; average recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 15.7s, FRENCH.
12709_13651_000760 · in -27.2 dBFS · gain +7.2 dB · mls-00054
This chain comes from the one-sided rule: only Shame had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Shame clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.36.
Nothing was asked of the other axis, and in fact Jealousy and Envy drifts down from 0.96 to 0.76 (-0.19), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.15, then +0.22, then -0.01 — not a clean run: step 3 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 59 s · polish · mls
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.948 before conversion and 0.930 after — it fell by 0.017. Neighbour-to-neighbour the worst pair went 0.935 → 0.898. (The earlier render, with segment 1 left raw, scores 0.794 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.356 in the original and +0.167 after conversion — 47 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Jealousy and Envy, -0.194 became -0.266.
Quality. Mean predicted overall quality across the segments went 3.10 → 3.41 (+0.31) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.948 → 0.930-0.017identity cos neighbours 0.935 → 0.898d_b rescored +0.356 → +0.167d_a rescored -0.194 → -0.266d_a mined -0.194d_b mined 0.356min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang polishspeaker 6892total 58.3schain gain +2.4 dBseam step 0.8 dBcrossfades 100/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert, slightly relaxed, fairly steady
(jealousy and envy, longing · normal-paced, average clarity, monologue, authoritative)jak prędko moglibyśmy się dostać do posiadłości lorda za sześć godzin jeśli weźmiemy pociąg odchodzący za trzy kwadranse w takim razie bądź pan gotów jedziemy natychmiast daruj pan ale nie rozumiem jeszcze w jakim celu mamy odbyć tę podróż
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, longing; style: monologue, authoritative; good recording, quiet background; genuineness 0.4/6; vocal-burst blend 0.6/10; 14.4s, POLISH.
6892_8338_000592 · in -26.1 dBFS · gain +6.1 dB · mls-00111
(sourness, disappointment·measured, clear, narration, monologue)poszukawszy w papierach wyjął notatkę jakąś i począł ją żywo przeglądać numer pięć pięć jest zapieczętowany własnym sygnetem mruczał nie nie mam prawa wyrzekł składając notatkę na dawne miejsce
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, disappointment; style: narration, monologue; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 2.6/10; 17.1s, POLISH.
6892_8338_000442 · in -26.6 dBFS · gain +6.5 dB · mls-00111
(shame, malevolence malice, disgust·normal-paced, average clarity, narration, authoritative)nie możesz pan w takim razie w imieniu lorda ja pana upoważniam bo od tego zależy życie pańskiego klienta jeśli łaska jedźmy natychmiast sir biggs spojrzał na mnie podejrzliwe z pod namarszczonych brwi
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, malevolence malice, disgust; style: narration, authoritative; good recording, quiet background; genuineness 0.4/6; vocal-burst blend 0.0/10; 13.0s, POLISH.
6892_8338_000712 · in -26.0 dBFS · gain +6.0 dB · mls-00111
(shame, disgust, malevolence malice · normal-paced, average clarity, monologue, narration)przepraszam ale mi się to wszystko coraz dziwniejszem wydaje wyrzekł po krótkiem milczeniu pozwól pan że ja z kolei zapytam czy masz pan dowody zarówno stwierdzające słowa pańskie jak upoważniające do rozkazywania mi w imieniu lorda puckinsa
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, disgust, malevolence malice; style: monologue, narration; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 1.1/10; 14.4s, POLISH.
6892_8338_000366 · in -25.2 dBFS · gain +5.2 dB · mls-00111
This chain comes from the one-sided rule: only Contempt had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Contempt clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.24.
Nothing was asked of the other axis, and in fact Malevolence Malice barely moves at all, sitting near 0.96 throughout.
It takes 5 clips to get there. Clip to clip the moves are +0.25, then -0.14, then +0.13, then -0.00 — not a clean run: step 2 moves back the other way by 0.14 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 101 s · portuguese · mls
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.897 before conversion and 0.877 after — it fell by 0.020. Neighbour-to-neighbour the worst pair went 0.933 → 0.906. (The earlier render, with segment 1 left raw, scores 0.743 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contempt moved +0.240 in the original and +0.132 after conversion — 55 % of the delta retained. On the other named axis, Malevolence Malice, +0.002 became +0.015.
Quality. Mean predicted overall quality across the segments went 2.93 → 3.45 (+0.52) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.897 → 0.877-0.020identity cos neighbours 0.933 → 0.906d_b rescored +0.240 → +0.132d_a rescored +0.002 → +0.015d_a mined 0.002d_b mined 0.241min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang portuguesespeaker 2961total 99.3schain gain +3.6 dBseam step 2.2 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an elderly somewhat feminine voice · slightly cool, rough, thin, quiet background, frequent disfluency, audible breath
(malevolence malice, distress, sadness · measured, subdued, slightly relaxed, whispered)na cova que está no campo de macpela que está em frente de manre na terra de canaã cova esta que abraão comprou de efrom o heteu juntamente com o respectivo campo como propriedade de sepultura ali sepultaram a abraão e a sara sua mulher ali sepultaram
full caption & clip details
An elderly somewhat feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, rough, thin; average clarity, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, distress, sadness; style: whispered, monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 2.6/10; 20.0s, PORTUGUESE.
2961_4205_000317 · in -23.7 dBFS · gain +3.7 dB · mls-00119
(malevolence malice, contempt, sourness·slow, very low-energy, relaxed, whispered)e a rebeca sua mulher e ali eu sepultei a léia o campo e a cova que está nele foram comprados aos filhos de hete acabando jacó de dar estas instruçães a seus filhos encolheu os seus pés na cama expirou e foi
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly cool, slightly dark, rough, thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is negative, neutral stance, slightly guarded; reads as malevolence malice, contempt, sourness; style: whispered, cartoonish; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 2.2/10; 20.0s, PORTUGUESE.
2961_4205_000281 · in -27.0 dBFS · gain +7.0 dB · mls-00119
(affection, contentment, thankfulness gratitude· slow, very low-energy, relaxed, cartoonish)ao seu povo então josé se lançou sobre o rosto de seu pai chorou sobre ele e o beijou e josé ordenou a seus servos os médicos que embalsamassem a seu pai e os médicos embalsamaram a israel
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, variable; timbre is slightly cool, dark, rough, thin; clear, frequent disfluency, narrow pitch range, audible breath; affect is negative, neutral stance, slightly guarded; reads as affection, contentment, thankfulness gratitude; style: cartoonish, monologue; below-average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.9/10; 20.0s, PORTUGUESE.
2961_4205_000293 · in -26.1 dBFS · gain +6.2 dB · mls-00119
(malevolence malice, disgust, sourness· slow, very low-energy, relaxed, cartoonish)olhos rogo-vos que faleis aos ouvidos de faraó dizendo meu pai me fez jurar dizendo eis que eu morro em meu sepulcro que cavei para mim na terra de canaã ali me sepultarás agora pois deixa-me subir peço-te e sepultar meu pai
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is slightly cool, slightly dark, rough, thin; clear, frequent disfluency, narrow pitch range, audible breath; affect is mildly positive, neutral stance, slightly guarded; reads as malevolence malice, disgust, sourness; style: cartoonish, whispered; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 1.6/10; 20.0s, PORTUGUESE.
2961_4205_000290 · in -26.7 dBFS · gain +6.7 dB · mls-00119
(contempt, malevolence malice, jealousy and envy· slow, subdued, slightly relaxed, cartoonish)dos egípcios pelo que o lugar foi chamado abel mizraim o qual está além do jordão assim os filhos de jacó lhe fizeram como ele lhes ordenara pois o levaram para a terra de canaã e o sepultaram na cova do campo de macpela que abraão tinha
full caption & clip details
An elderly somewhat feminine voice; delivery is subdued, slow, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, rough, thin; average clarity, frequent disfluency, narrow pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, malevolence malice, jealousy and envy; style: cartoonish, whispered; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 1.3/10; 20.0s, PORTUGUESE.
2961_4205_000282 · in -25.4 dBFS · gain +5.4 dB · mls-00119
This chain comes from the one-sided rule: only Pain had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Pain clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.27.
Nothing was asked of the other axis, and in fact Contempt drifts down from 0.99 to 0.77 (-0.22), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.18, then +0.08, then +0.01 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 65 s · french · mls
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.935 before conversion and 0.861 after — it fell by 0.073. Neighbour-to-neighbour the worst pair went 0.935 → 0.861. (The earlier render, with segment 1 left raw, scores 0.811 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.448 in the original and +0.633 after conversion — 141 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contempt, -0.222 became -0.051.
Quality. Mean predicted overall quality across the segments went 3.11 → 3.26 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.935 → 0.861-0.073identity cos neighbours 0.935 → 0.861d_b rescored +0.448 → +0.633d_a rescored -0.222 → -0.051d_a mined -0.222d_b mined 0.275min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang frenchspeaker 12709total 63.7schain gain +1.4 dBseam step 0.9 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, clear
(contempt, anger, sourness · measured, almost no disfluency, moderate pitch range, monologue)on vous criera aux oreilles les intérêts des particuliers ne sont rien en concurrence avec l'intérêt du tout combien il est facile d'avancer une maxime générale que personne n'ose contester
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, anger, sourness; style: monologue, narration; good recording, quiet background; genuineness 0.7/6; vocal-burst blend 1.1/10; 12.7s, FRENCH.
12709_14409_000106 · in -27.1 dBFS · gain +7.1 dB · mls-00058
(pride, triumph, relief·normal-paced, no disfluency, moderate pitch range, narration)mais qu'il est difficile et rare d'avoir toutes les connaissances de détail nécessaires pour en prévenir une fausse application heureusement pour moi monsieur et pour vous j'ai à peu près exercé la double profession d'auteur et de libraire j'ai écrit et j'ai plusieurs fois imprimé pour mon compte
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, triumph, relief; style: narration, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 1.5/10; 19.2s, FRENCH.
12709_14409_000474 · in -27.2 dBFS · gain +7.2 dB · mls-00058
(emotional numbness, bitterness, malevolence malice· normal-paced, no disfluency, moderate pitch range, narration)et je puis vous assurer chemin faisant que rien ne s'accorde plus mal que la vie active du commerçant et la vie sédentaire de l'homme de lettres incapables que nous sommes d'une infinité de petits soins sur cent auteurs qui voudront débiter eux-mêmes leurs ouvrages
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, bitterness, malevolence malice; style: narration, monologue; good recording, quiet background; genuineness 0.1/6; vocal-burst blend 2.1/10; 15.5s, FRENCH.
12709_14409_000636 · in -26.8 dBFS · gain +6.8 dB · mls-00058
(pain, sadness, distress·measured, almost no disfluency, fairly narrow pitch, whispered)il y en a quatre-vingt-dix-neuf qui s'en trouveront mal et s'en dégoûteront le libraire peu scrupuleux croit que l'auteur court sur ses brisées lui qui jette les hauts cris quand on le contrefait qui se tiendrait pour malhonnête homme s'il contrefaisait son confrère
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain, sadness, distress; style: whispered, didactic; average recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.9/10; 16.9s, FRENCH.
12709_14409_000096 · in -27.5 dBFS · gain +7.5 dB · mls-00058
This chain comes from the one-sided rule: only Pain had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Pain strongly present — 0.80, higher than 80 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Contentment drifts down from 1.00 to 0.24 (-0.76), which the rule did not require.
It takes 2 clips to get there. Clip to clip the moves are +0.20 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 33 s · italian · mls
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.930 before conversion and 0.938 after — it rose by 0.007. Neighbour-to-neighbour the worst pair went 0.930 → 0.938. (The earlier render, with segment 1 left raw, scores 0.873 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.282 in the original and +0.282 after conversion — 100 % of the delta retained, which is essentially all of it. On the other named axis, Contentment, -0.764 became -0.763.
Quality. Mean predicted overall quality across the segments went 3.13 → 3.33 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.930 → 0.938+0.007identity cos neighbours 0.930 → 0.938d_b rescored +0.282 → +0.282d_a rescored -0.764 → -0.763d_a mined -0.764d_b mined 0.202min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang italianspeaker 12804total 32.9schain gain -1.2 dBseam step 0.6 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an elderly masculine voice · slightly dark, balanced body, no background noise, slightly relaxed, steady, little disfluency, somewhat unclear, fairly narrow pitch
(contentment, thankfulness gratitude, shame · measured, subdued, monologue, narration)la qual presta a comandamenti della signora in tal maniera disse vien da le parti di settentrione gente rubesta di bianco vestita ferisse ogn un senza compassione nel capo ne li piedi e ne la vita
full caption & clip details
An elderly masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, little disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as contentment, thankfulness gratitude, shame; style: monologue, narration; average recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.9/10; 18.8s, ITALIAN.
12804_11418_000071 · in -29.0 dBFS · gain +9.0 dB · mls-00066
(pain, distress, sadness·slow, very low-energy, monologue, narration)di morti stan coperte le persone e di salvarsi ogn un qua e là s'aita arde in le case d ogni canto il fuoco da lor schermirsi non si trova luoco
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is slightly warm, slightly dark, rough, balanced body; somewhat unclear, little disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as pain, distress, sadness; style: monologue, narration; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 1.7/10; 14.4s, ITALIAN.
12804_11418_000110 · in -28.1 dBFS · gain +8.1 dB · mls-00066