proxy_spearman__PXR__T0.20__C0.25__INTERNAL — voice-corrected

Manifest tier. proxy_spearman, rule PXR, T=0.2, step cap 0.25. Population 1,222,387 chains (12,216 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 958,511.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_proxy_spearman__PXR__T0.20__C0.25__INTERNAL.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
44segments re-voiced
0.774 → 0.802median worst-to-anchor identity cosine
90 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Shame ↓  /  Contemptidentity −0.06 emotion 35 %   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #1

This chain comes from the proxy rule: the same two-sided test as above, but because Contempt is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contempt strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.21.

At the same time Shame goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.77 (higher than 77 % of clips in this corpus), a change of -0.22. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are -0.06, then +0.10, then +0.16 — not a clean run: step 1 moves back the other way by 0.06 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 65 s · it · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.919 before conversion and 0.856 after — it fell by 0.063. Neighbour-to-neighbour the worst pair went 0.929 → 0.815. (The earlier render, with segment 1 left raw, scores 0.705 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contempt moved +0.205 in the original and +0.072 after conversion — 35 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Shame, -0.223 became -0.126.

Quality. Mean predicted overall quality across the segments went 2.98 → 3.38 (+0.40) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.919 → 0.856 -0.063identity cos neighbours 0.929 → 0.815d_b rescored +0.205 → +0.072d_a rescored -0.223 → -0.126d_a mined -0.223d_b mined 0.206min_cos_consec (site) —min_cos_anchor (site) —dataset eurospeechlang itspeaker italy_15_112total 64.5schain gain +2.6 dBseam step 1.2 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · quiet background, moderately variable
(shame, concentration, thankfulness gratitude · brisk, energised, neutral tension, authoritative) che la NATO si isolasse, facendo della missione afghana una sfida solo della NATO. La missione afghana è innanzi tutto una sfida dell'intera comunità internazionale,
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as shame, concentration, thankfulness gratitude; style: authoritative, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 3.7/10; 15.7s, IT.
italy_15_112_3123551_3139232 · in -21.9 dBFS · gain +1.9 dB · eurospeech-01636
(shame, jealousy and envy, contentment · normal-paced, normally alert, neutral tension, authoritative) delle Nazioni Unite e dell'insieme dei Paesi del mondo, tra i quali - faccio osservare - non ve n'è neppure uno che sostenga la necessità di ritirare le forze internazionali dall'Afghanistan,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as shame, jealousy and envy, contentment; style: authoritative, monologue; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 6.1/10; 15.5s, IT.
italy_15_112_3139232_3154703 · in -21.4 dBFS · gain +1.4 dB · eurospeech-01636
(impatience and irritability, anger, shame · brisk, energised, neutral tension, authoritative) dal momento che tutti i Paesi del mondo - tra i quali ne cito due piuttosto importanti nella regione: la Russia e la Cina - ritengono che un ritorno dei talibani sarebbe una tragedia
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, anger, shame; style: authoritative, casual; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 3.7/10; 16.1s, IT.
italy_15_112_3154703_3170783 · in -21.7 dBFS · gain +1.7 dB · eurospeech-01636
(contempt, anger, bitterness · measured, energised, tense, cartoonish) non accettabile, anche per loro. La Cina ha 93 chilometri di confine con l'Afghanistan. È dunque necessario impegnare l'insieme di questi Paesi in uno sforzo comune.
full caption & clip details
A middle-aged masculine voice; delivery is energised, measured, tense, moderately variable; timbre is slightly cool, neutral-bright, rough, thin; average clarity, frequent disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, fairly guarded; reads as contempt, anger, bitterness; style: cartoonish, authoritative; below-average recording, quiet background; genuineness 2.4/6; vocal-burst blend 2.6/10; 17.8s, IT.
italy_15_112_3170783_3188592 · in -20.7 dBFS · gain +0.7 dB · eurospeech-01636
Emotional Numbness ↓  /  Triumphidentity −0.03 emotion 135 %   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #2

This chain comes from the proxy rule: the same two-sided test as above, but because Triumph is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Triumph strongly present — 0.77, higher than 77 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.23.

At the same time Emotional Numbness goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.05, then +0.18 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 45 s · dutch · mls

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.895 before conversion and 0.864 after — it fell by 0.030. Neighbour-to-neighbour the worst pair went 0.869 → 0.826. (The earlier render, with segment 1 left raw, scores 0.754 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Triumph moved +0.226 in the original and +0.307 after conversion — 135 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, -0.281 became -0.203.

Quality. Mean predicted overall quality across the segments went 3.28 → 3.47 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.895 → 0.864 -0.030identity cos neighbours 0.869 → 0.826d_b rescored +0.226 → +0.307d_a rescored -0.281 → -0.203d_a mined -0.281d_b mined 0.225min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang nlspeaker 2450total 44.1schain gain +2.6 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, slightly dark, rough, no background noise, slightly relaxed, steady, almost no disfluency, fairly narrow pitch
(emotional numbness · slow, normally alert, very clear, narration) en gerimpelde huid in hooge mate vleiend was beloven doe ik niets maar ik zal zien schikt het u als we van avond een visite brengen
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, rough, balanced body; very clear, almost no disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: narration, storytelling; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.2/10; 12.6s, DUTCH.
2450_12965_000633 · in -28.1 dBFS · gain +8.1 dB · mls-00091
(intoxication altered states of consciousness, sexual lust, thankfulness gratitude · slow, normally alert, slurred, narration) met genoegen wel zeker hij werd geroepen bij den vriend die hij kwam bep a daum hoe hij raad van indië werd onder ps maurits zoeken bij het heengaan drukte hij louise's hand teeder en zag haar veelbeteekenend aan
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, rough, thin; slurred, almost no disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness, sexual lust, thankfulness gratitude; style: narration, storytelling; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.2/10; 12.6s, DUTCH.
2450_12965_000800 · in -28.5 dBFS · gain +8.5 dB · mls-00091
(triumph, anger, concentration · measured, subdued, slurred, narration) wel vroeg ze aan kees die een tien minuten later thuis kwam met de oude vrouw is alles in orde wij kunnen er komen dat is heerlijk laat de rest maar aan mij over den ouden heb ik niet kunnen spreken hij was niet op zijn bureau
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, rough, balanced body; slurred, almost no disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, anger, concentration; style: narration, storytelling; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.3/10; 19.3s, DUTCH.
2450_12965_000388 · in -27.1 dBFS · gain +7.1 dB · mls-00091
Concentration ↓  /  Fearidentity −0.02 emotion 80 %   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #3

This chain comes from the proxy rule: the same two-sided test as above, but because Fear is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fear clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.27.

At the same time Concentration goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.09, then +0.18 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 30 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.797 before conversion and 0.774 after — it fell by 0.023. Neighbour-to-neighbour the worst pair went 0.797 → 0.774. (The earlier render, with segment 1 left raw, scores 0.711 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.268 in the original and +0.215 after conversion — 80 % of the delta retained, which is most of it. On the other named axis, Concentration, -0.357 became -0.570.

Quality. Mean predicted overall quality across the segments went 2.79 → 2.96 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.797 → 0.774 -0.023identity cos neighbours 0.797 → 0.774d_b rescored +0.268 → +0.215d_a rescored -0.357 → -0.570d_a mined -0.357d_b mined 0.268min_cos_consec (site) 0.9304min_cos_anchor (site) 0.9387dataset emolialang enspeaker EN_OLC11JRdtuktotal 29.0schain gain +1.2 dBseam step 0.3 dBcrossfades 100/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · quiet background, normally alert, moderate pitch range
(concentration, pain · normal-paced, neutral tension, moderately variable, casual) That gives you the, the, the, the dye yield is a, a portion process and that, that's e raised to the power minus lambda, where lambda is the average number of defects that may occur per dye. And that number of defects is equal to the defect density and dye area.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration, pain; style: casual, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 4.2/10; 16.5s, EN.
EN_OLC11JRdtuk_W000189 · in -18.6 dBFS · gain -1.4 dB · emolia-01356
(measured, slightly relaxed, fairly steady, didactic) Defect density is (low mumble) somewhere between 0.2 to 1 per centimeter square.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, formal; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 0.2/10; 6.0s, EN.
EN_OLC11JRdtuk_W000190 · in -16.8 dBFS · gain -3.2 dB · emolia-01356
(measured, slightly relaxed, moderately variable, casual) If you look at the, the, the manufactured device (low mumble) right from 60s or 70s.
full caption & clip details
A child masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is slightly cool, dark, slightly rough, thin; average clarity, frequent disfluency, moderate pitch range, audible breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: casual, playful; below-average recording, quiet background; genuineness 3.3/6; vocal-burst blend 3.4/10; 6.8s, EN.
EN_OLC11JRdtuk_W000191 · in -18.6 dBFS · gain -1.4 dB · emolia-01356
Interest ↓  /  Affectionidentity −0.07 emotion 113 %   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #4

This chain comes from the proxy rule: the same two-sided test as above, but because Affection is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Affection clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.31.

At the same time Interest goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.18, then +0.13 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 44 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.874 before conversion and 0.808 after — it fell by 0.066. Neighbour-to-neighbour the worst pair went 0.874 → 0.808. (The earlier render, with segment 1 left raw, scores 0.585 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Affection moved +0.313 in the original and +0.353 after conversion — 113 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Interest, -0.262 became -0.337.

Quality. Mean predicted overall quality across the segments went 2.82 → 3.27 (+0.45) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.874 → 0.808 -0.066identity cos neighbours 0.874 → 0.808d_b rescored +0.313 → +0.353d_a rescored -0.262 → -0.337d_a mined -0.262d_b mined 0.313min_cos_consec (site) 0.8533min_cos_anchor (site) 0.8533dataset emolialang enspeaker EN_PFWjw2nMG2Ytotal 42.9schain gain +5.1 dBseam step 0.8 dBcrossfades 100/100 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, slightly bright, fairly smooth, quiet background, normal-paced, normally alert, slightly relaxed, fairly steady
(interest, elation · frequent disfluency, normal breath, casual, monologue) Yeah, so (ahem) uhm, so there was an article about inclusion in the workplace, (ahem) uhm, some metrics, (ahem) uh, it was going around, and it was like the top, I don't know, five metrics of, (ahem) uhm, what was, (breathy giggle) I guess, important or crucial to this feeling of inclusion in the workplace.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as interest, elation; style: casual, monologue; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.0/10; 18.0s, EN.
EN_PFWjw2nMG2Y_W000015 · in -17.4 dBFS · gain -2.6 dB · emolia-00679
(thankfulness gratitude · frequent disfluency, light breath, casual, monologue) So (childlike giggle) uhm, I looked at those and there were a few that we did not have metrics around. So (ahem) uhm, that's where those ideas came from. There was (ahem) uhm, I think psychological safety, (low mumble) uhm, can't remember what the other ones were, uh, (low mumble) but,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as thankfulness gratitude; style: casual, monologue; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 3.0/10; 13.7s, EN.
EN_PFWjw2nMG2Y_W000016 · in -19.5 dBFS · gain -0.5 dB · emolia-00679
(affection · some disfluency, light breath, casual, monologue) Fair, oh, fairness and (ahem) psychological support. So (ahem) uhm, I just took the trust and safety one, (wistful sigh) uh, and just started developing a metric from scratch based on it.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as affection; style: casual, monologue; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 1.8/10; 11.6s, EN.
EN_PFWjw2nMG2Y_W000017 · in -20.2 dBFS · gain +0.2 dB · emolia-00679
Astonishment Surprise ↓  /  Hope Enthusiasm Optimismidentity +0.01 emotion 82 %   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #5

This chain comes from the proxy rule: the same two-sided test as above, but because Hope Enthusiasm Optimism is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Hope Enthusiasm Optimism around average — 0.57, higher than 57 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.42.

At the same time Astonishment Surprise goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.48 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.47 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.48, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 18 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.523 before conversion and 0.532 after — it rose by 0.010. Neighbour-to-neighbour the worst pair went 0.473 → 0.532. (The earlier render, with segment 1 left raw, scores 0.472 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.417 in the original and +0.341 after conversion — 82 % of the delta retained, which is most of it. On the other named axis, Astonishment Surprise, -0.339 became -0.637.

Quality. Mean predicted overall quality across the segments went 2.68 → 2.90 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.523 → 0.532 +0.010identity cos neighbours 0.473 → 0.532d_b rescored +0.417 → +0.341d_a rescored -0.339 → -0.637d_a mined -0.339d_b mined 0.417min_cos_consec (site) 0.4703min_cos_anchor (site) 0.4756dataset emolialang enspeaker EN_GKkP_YE3D4ytotal 17.1schain gain +2.6 dBseam step 0.5 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, quiet background, normal-paced, normally alert, some disfluency
(astonishment surprise, impatience and irritability, anger · slightly relaxed, fairly steady, average clarity, casual) (low mumble) Uhm, this was a fucking insane video. We're gonna spin down to 17k.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as astonishment surprise, impatience and irritability, anger; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 3.5/6; vocal-burst blend 0.6/10; 5.2s, EN.
EN_GKkP_YE3D4y_W000148 · in -19.1 dBFS · gain -0.9 dB · emolia-01974
(sexual lust, pleasure ecstasy, intoxication altered states of consciousness · fully relaxed, moderately variable, slurred, casual) What a video, dude. Oh my, can't get over this.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, moderately variable; timbre is neutral-toned, dark, slightly rough, thin; slurred, some disfluency, moderate pitch range, audible breath; affect is mildly negative, neutral stance, neutral openness; reads as sexual lust, pleasure ecstasy, intoxication altered states of consciousness; style: casual, conversational; below-average recording, quiet background; mildly explicit content; genuineness 5.0/6; vocal-burst blend 1.1/10; 3.5s, EN.
EN_GKkP_YE3D4y_W000151 · in -19.6 dBFS · gain -0.4 dB · emolia-01974
(hope enthusiasm optimism, contentment, elation · neutral tension, moderately variable, average clarity, casual) Alright, and that is that. I hope y'all enjoyed the video. I'll see y'all tomorrow with another one and probably another video in here in a couple of days. Peace out, guys.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism, contentment, elation; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 8.5/10; 8.8s, EN.
EN_GKkP_YE3D4y_W000152 · in -23.0 dBFS · gain +3.0 dB · emolia-01974
Concentration ↓  /  Intoxication Altered States of Consciousnessidentity −0.08 emotion 176 %   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #6

This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Intoxication Altered States of Consciousness around average — 0.49, right about the corpus median — and ends with it clearly present at 0.74, higher than 74 % of clips in this corpus. That is a total rise of 0.25.

At the same time Concentration goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.58 (higher than 58 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.07, then +0.16, then +0.01 — a plateau around step 3, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 26 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.799 before conversion and 0.715 after — it fell by 0.085. Neighbour-to-neighbour the worst pair went 0.813 → 0.715. (The earlier render, with segment 1 left raw, scores 0.617 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.246 in the original and +0.432 after conversion — 176 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.281 became -0.297.

Quality. Mean predicted overall quality across the segments went 2.93 → 2.99 (+0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.799 → 0.715 -0.085identity cos neighbours 0.813 → 0.715d_b rescored +0.246 → +0.432d_a rescored -0.281 → -0.297d_a mined -0.280d_b mined 0.246min_cos_consec (site) 0.8146min_cos_anchor (site) 0.8192dataset emolialang enspeaker EN_B00058_S06573total 24.6schain gain +1.2 dBseam step 2.1 dBcrossfades 150/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, moderate pitch range, light breath
(normal-paced, fairly steady, some disfluency, casual) One of the easiest ways is to simply create a variable in our calculate view controller. So we can call it BMI value.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, dramatic; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 0.8/10; 7.8s, EN.
EN_B00058_S06573_W000105 · in -19.7 dBFS · gain -0.3 dB · emolia-01359
(normal-paced, moderately variable, some disfluency, whispered) And we can set it to start off being equal to 0.0. And then when we calculate our BMI,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: whispered, casual; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 1.0/10; 6.5s, EN.
EN_B00058_S06573_W000106 · in -19.9 dBFS · gain -0.1 dB · emolia-01359
(measured, moderately variable, some disfluency, casual) We can format it so that we only get it to one decimal place. So we could say BMI value.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 3.1/10; 6.4s, EN.
EN_B00058_S06573_W000107 · in -20.2 dBFS · gain +0.2 dB · emolia-01359
(normal-paced, fairly steady, little disfluency, whispered) With this particular format. So it's going to be %.1f.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: whispered, casual; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 0.8/10; 4.5s, EN.
EN_B00058_S06573_W000108 · in -21.6 dBFS · gain +1.6 dB · emolia-01359
Longing ↓  /  Impatience and Irritabilityidentity +0.24 emotion REVERSED   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #7

This chain comes from the proxy rule: the same two-sided test as above, but because Impatience and Irritability is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Impatience and Irritability clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.26.

At the same time Longing goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.67 (higher than 67 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.01 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.53 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.57 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.53, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 16 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.507 before conversion and 0.749 after — it rose by 0.242. Neighbour-to-neighbour the worst pair went 0.590 → 0.810. (The earlier render, with segment 1 left raw, scores 0.697 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Impatience and Irritability moved +0.262 in the original and -0.064 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Longing, -0.322 became -0.286.

Quality. Mean predicted overall quality across the segments went 2.54 → 2.81 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.507 → 0.749 +0.242identity cos neighbours 0.590 → 0.810d_b rescored +0.262 → -0.064d_a rescored -0.322 → -0.286d_a mined -0.322d_b mined 0.261min_cos_consec (site) 0.5723min_cos_anchor (site) 0.5323dataset podcastlang enspeaker 604898total 15.4schain gain +3.9 dBseam step 1.5 dBcrossfades 100/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · thin, average recording, quiet background, moderately variable, average clarity
(longing, fatigue exhaustion, contemplation · normal-paced, normally alert, relaxed, casual) so many years you know we've been this way and
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, dark, slightly rough, thin; average clarity, frequent disfluency, moderate pitch range, audible breath; affect is mildly positive, slightly submissive, neutral openness; reads as longing, fatigue exhaustion, contemplation; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 5.5/10; 3.5s, EN.
604898_00074271 · in -31.4 dBFS · gain +11.4 dB · podcast-03132
(jealousy and envy, disgust, bitterness · brisk, energised, neutral tension, casual) and then the closeness that you get with him and the faith that you build because it's like okay
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as jealousy and envy, disgust, bitterness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 3.8/6; vocal-burst blend 3.8/10; 4.7s, EN.
604898_00078272 · in -23.5 dBFS · gain +3.5 dB · podcast-03111
(impatience and irritability, bitterness, triumph · brisk, energised, neutral tension, casual) not the end of like you said the prize is not that it's everything that you encounter that journey the roadblocks the detours that's the process.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as impatience and irritability, bitterness, triumph; style: casual, dramatic; average recording, quiet background; mildly explicit content; genuineness 4.3/6; vocal-burst blend 6.3/10; 7.5s, EN.
604898_00079216 · in -24.4 dBFS · gain +4.4 dB · podcast-03121
Doubt ↓  /  Fatigue Exhaustionidentity +0.05 emotion 55 %   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #8

This chain comes from the proxy rule: the same two-sided test as above, but because Fatigue Exhaustion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fatigue Exhaustion clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.29.

At the same time Doubt goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.23. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.17, then +0.12, then +0.01 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.75 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.75 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 20 s · ja · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.654 before conversion and 0.704 after — it rose by 0.050. Neighbour-to-neighbour the worst pair went 0.665 → 0.704. (The earlier render, with segment 1 left raw, scores 0.706 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.318 in the original and +0.175 after conversion — 55 % of the delta retained. On the other named axis, Doubt, -0.230 became -0.042.

Quality. Mean predicted overall quality across the segments went 3.00 → 3.01 (+0.02) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.654 → 0.704 +0.050identity cos neighbours 0.665 → 0.704d_b rescored +0.318 → +0.175d_a rescored -0.230 → -0.042d_a mined -0.230d_b mined 0.292min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang jaspeaker JA_B00002_S01872total 19.0schain gain -2.0 dBseam step 1.1 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, slightly relaxed, fairly steady, average clarity
(doubt, confusion · fast, normally alert, no disfluency, storytelling) 明日の歯医者の予約時間、手帳に書いとくべきだったなぁ。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, confusion; style: storytelling, narration; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 6.0/10; 3.7s, JA.
JA_B00002_S01872_W000028 · in -17.7 dBFS · gain -2.3 dB · emolia-02967
(measured, subdued, frequent disfluency, conversational) (resonant hum) うーん、自分に合う道具を使うことかな。
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational, monologue; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 2.9/10; 3.0s, JA.
JA_B00002_S01872_W000029 · in -16.4 dBFS · gain -3.6 dB · emolia-02967
(teasing, jealousy and envy, sexual lust · measured, normally alert, no disfluency, conversational) (resonant hum) 上達しにくいものは特にないよ。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as teasing, jealousy and envy, sexual lust; style: conversational, storytelling; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 4.3/10; 3.2s, JA.
JA_B00002_S01872_W000030 · in -18.2 dBFS · gain -1.8 dB · emolia-02967
(fatigue exhaustion, contemplation · measured, normally alert, some disfluency, monologue) この店は、社会人や御高齢のお客様の割合が多いので、そういった方にゆっくりしてもらえるよう、緑あふれる明るい庭を作ってはどうですか?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, contemplation; style: monologue, authoritative; average recording, no background noise; genuineness 1.8/6; vocal-burst blend 5.5/10; 9.8s, JA.
JA_B00002_S01872_W000031 · in -16.6 dBFS · gain -3.4 dB · emolia-02967
Sexual Lust ↓  /  Astonishment Surpriseidentity −0.07 emotion 100 %   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #9

This chain comes from the proxy rule: the same two-sided test as above, but because Astonishment Surprise is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Astonishment Surprise clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.33.

At the same time Sexual Lust goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.74 (higher than 74 % of clips in this corpus), a change of -0.21. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.10, then +0.24 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 39 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.634 before conversion and 0.565 after — it fell by 0.069. Neighbour-to-neighbour the worst pair went 0.634 → 0.565. (The earlier render, with segment 1 left raw, scores 0.604 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.335 in the original and +0.334 after conversion — 100 % of the delta retained, which is essentially all of it. On the other named axis, Sexual Lust, -0.218 became -0.411.

Quality. Mean predicted overall quality across the segments went 2.81 → 3.03 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.634 → 0.565 -0.069identity cos neighbours 0.634 → 0.565d_b rescored +0.335 → +0.334d_a rescored -0.218 → -0.411d_a mined -0.211d_b mined 0.335min_cos_consec (site) 0.8547min_cos_anchor (site) 0.8547dataset podcastlang enspeaker 273819total 38.6schain gain +5.0 dBseam step 3.0 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, quiet background
(sexual lust, teasing · normal-paced, normally alert, fully relaxed, casual) said, I really listened to it. Okay, let's go on this tangent real quick, okay?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, thin; slurred, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as sexual lust, teasing; style: casual, conversational; poor recording, quiet background; mildly explicit content; genuineness 5.2/6; vocal-burst blend 2.2/10; 3.6s, EN.
273819_00164560 · in -29.7 dBFS · gain +9.7 dB · podcast-04011
(intoxication altered states of consciousness, fatigue exhaustion, elation · normal-paced, normally alert, neutral tension, casual) at the mall. We have this anime store at the mall called Anime GT. And I'm about to walk in there. I used to work at the mall, so I I used to just you know head in there if I can look around a little bit. I w I'm walking, I'm literally about to walk in. This is like one of the first times I'm about to walk in by myself. I wasn't that into anime at the time. But I I knew of anime and I watched some.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as intoxication altered states of consciousness, fatigue exhaustion, elation; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.9/6; vocal-burst blend 10.0/10; 22.4s, EN.
273819_00165048 · in -31.2 dBFS · gain +11.2 dB · podcast-02801
(astonishment surprise, anger, sourness · measured, energised, neutral tension, casual) And I'm about to walk in, and the this is the only anime that I heard was terrible was Sword Art Online. Like ass after the first season. And then
full caption & clip details
A young adult masculine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as astonishment surprise, anger, sourness; style: casual, monologue; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.6/10; 13.0s, EN.
273819_00167288 · in -31.8 dBFS · gain +11.8 dB · podcast-02806
Confusion ↓  /  Interestidentity +0.10 emotion 143 %   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #10

This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Interest clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.22.

At the same time Confusion goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.58 (higher than 58 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.14, then +0.08 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.87 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.87 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 51 s · ko · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.743 before conversion and 0.845 after — it rose by 0.102. Neighbour-to-neighbour the worst pair went 0.829 → 0.845. (The earlier render, with segment 1 left raw, scores 0.636 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.218 in the original and +0.313 after conversion — 143 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Confusion, -0.426 became -0.168.

Quality. Mean predicted overall quality across the segments went 2.79 → 3.19 (+0.39) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.743 → 0.845 +0.102identity cos neighbours 0.829 → 0.845d_b rescored +0.218 → +0.313d_a rescored -0.426 → -0.168d_a mined -0.332d_b mined 0.218min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang kospeaker KO_ZPIJVlLlS8ctotal 49.9schain gain -1.9 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(confusion · monologue, conversational) 그러니까 저는 지금 이제 사실 뭐라 그래도 지금 제주에, 제도도의 인근은 제주시 지역에, 동지역에 몰려있지 않습니까? 많이 집중되어 있죠. 어, 죽어도 그렇고, 활동도 그렇고.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion; style: monologue, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 6.0/10; 12.3s, KO.
KO_ZPIJVlLlS8c_W000020 · in -13.9 dBFS · gain -6.1 dB · emolia-03176
(contemplation, jealousy and envy, disappointment · monologue, conversational) 그래서 사실 그런 측면에서 보면은, 이제, 읍면 지역에 인구가 그렇게 많지 않기 때문에, 앞으로 제주 지역을, 글쎄요, 제주시 지역만 집중발류하고, 저, 개발을 하고, 읍면 지역을 지금처럼, 어, (resonant hum) 반동반호, 소위 전통적인 얘기에, 거기에 약간 관광이 가미된, 그런 정도 수준으로 하겠다면,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, jealousy and envy, disappointment; style: monologue, conversational; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 8.0/10; 19.5s, KO.
KO_ZPIJVlLlS8c_W000021 · in -13.7 dBFS · gain -6.3 dB · emolia-03176
(interest, anger, astonishment surprise · monologue, conversational) 어느 정도, 이렇게, 이렇게, 이 말이 뭐 정확하게 좀 약간 어패가 좀 있을 수 있는데 도시화, 도시화가 되면 인구는 불안하게 되어 있거든요. 왜냐하면 아무리 농춘을 우리가 전원적인 삶을 주행한다 그래도 기반시설이 약하기 때문에 인구 유입이 사실 어려워요. 그런데 도시화된 지역에는, 거기는 당연히 학교도 있고,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, anger, astonishment surprise; style: monologue, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 9.8/10; 18.4s, KO.
KO_ZPIJVlLlS8c_W000022 · in -13.0 dBFS · gain -7.0 dB · emolia-03176
Concentration ↓  /  Astonishment Surpriseidentity +0.02 emotion 96 %   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #11

This chain comes from the proxy rule: the same two-sided test as above, but because Astonishment Surprise is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Astonishment Surprise clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.23.

At the same time Concentration goes the other way, from 0.77 (higher than 77 % of clips in this corpus) to 0.55 (higher than 55 % of clips in this corpus), a change of -0.22. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 20 s · snippets

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.863 before conversion and 0.884 after — it rose by 0.021. Neighbour-to-neighbour the worst pair went 0.863 → 0.884. (The earlier render, with segment 1 left raw, scores 0.786 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.231 in the original and +0.223 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.223 became -0.138.

Quality. Mean predicted overall quality across the segments went 3.08 → 3.16 (+0.08) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.863 → 0.884 +0.021identity cos neighbours 0.863 → 0.884d_b rescored +0.231 → +0.223d_a rescored -0.223 → -0.138d_a mined -0.219d_b mined 0.231min_cos_consec (site) —min_cos_anchor (site) —dataset snippetslang undspeaker batch248_part2_batch248_patotal 19.7schain gain +3.3 dBseam step 0.6 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normal-paced, normally alert, slightly relaxed
(almost no disfluency, narration, formal) an allies the attackers along the existing road. Brodka had been smart and understanding the importance of the two towers on the eastern end of the bridge.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.1/10; 8.6s.
batch248_part2_batch248_part2_chunk_702_1_638157 · in -27.3 dBFS · gain +7.3 dB · snippets-00773
(astonishment surprise, awe, disappointment · some disfluency, dramatic, monologue) And the troops were exhausted. Many had partaken in the battle of the bulge and had marched almost non-stop since then. A truly fatigued soldier is never a great thing on the battlefield.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as astonishment surprise, awe, disappointment; style: dramatic, monologue; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.3/10; 11.3s.
batch248_part2_batch248_part2_chunk_702_1_638189 · in -27.6 dBFS · gain +7.6 dB · snippets-00773
Fatigue Exhaustion ↓  /  Painidentity +0.04 emotion 41 %   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #12

This chain comes from the proxy rule: the same two-sided test as above, but because Pain is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Pain around average — 0.54, higher than 54 % of clips in this corpus — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.29.

At the same time Fatigue Exhaustion goes the other way, from 0.77 (higher than 77 % of clips in this corpus) to 0.47 (lower than 53 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.17, then +0.12 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 21 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.820 before conversion and 0.855 after — it rose by 0.035. Neighbour-to-neighbour the worst pair went 0.811 → 0.851. (The earlier render, with segment 1 left raw, scores 0.796 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.292 in the original and +0.119 after conversion — 41 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Fatigue Exhaustion, -0.296 became -0.184.

Quality. Mean predicted overall quality across the segments went 2.90 → 3.16 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.820 → 0.855 +0.035identity cos neighbours 0.811 → 0.851d_b rescored +0.292 → +0.119d_a rescored -0.296 → -0.184d_a mined -0.296d_b mined 0.292min_cos_consec (site) 0.8455min_cos_anchor (site) 0.8408dataset emolialang zhspeaker ZH_B00029_S06915total 20.2schain gain +2.2 dBseam step 2.2 dBcrossfades 100/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(normal-paced, fairly steady, no disfluency, formal) 在南进书区找到一本王形龙的楷贴。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, didactic; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.6/10; 4.0s, ZH.
ZH_B00029_S06915_W000094 · in -19.1 dBFS · gain -0.9 dB · emolia-03567
(measured, fairly steady, no disfluency, formal) 宁缺抽出来,一面研读,一面随意行走。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.9/10; 4.1s, ZH.
ZH_B00029_S06915_W000095 · in -19.5 dBFS · gain -0.5 dB · emolia-03567
(measured, steady, no disfluency, formal) 渐渐的,身旁变得越来越安静,他抬起头来。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, didactic; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.1/10; 5.5s, ZH.
ZH_B00029_S06915_W000096 · in -19.0 dBFS · gain -1.0 dB · emolia-03567
(measured, fairly steady, almost no disfluency, didactic) 只见一道干净的楼梯出现在眼前,楼梯是用来上楼的。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 1.3/10; 7.2s, ZH.
ZH_B00029_S06915_W000097 · in -18.1 dBFS · gain -1.9 dB · emolia-03567
Interest ↓  /  Confusionidentity −0.01 emotion 100 %   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #13

This chain comes from the proxy rule: the same two-sided test as above, but because Confusion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Confusion around average — 0.49, lower than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.48.

At the same time Interest goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.79 (higher than 79 % of clips in this corpus), a change of -0.21. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then -0.03, then +0.20, then +0.09 — not a clean run: step 2 moves back the other way by 0.03 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.94 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.94 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 64 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.858 before conversion and 0.850 after — it fell by 0.008. Neighbour-to-neighbour the worst pair went 0.839 → 0.848. (The earlier render, with segment 1 left raw, scores 0.743 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.482 in the original and +0.480 after conversion — 100 % of the delta retained, which is essentially all of it. On the other named axis, Interest, -0.208 became -0.232.

Quality. Mean predicted overall quality across the segments went 2.86 → 3.01 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.858 → 0.850 -0.008identity cos neighbours 0.839 → 0.848d_b rescored +0.482 → +0.480d_a rescored -0.208 → -0.232d_a mined -0.207d_b mined 0.482min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_JPwc5UCV-bototal 62.3schain gain +2.4 dBseam step 1.0 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a child feminine voice · slightly thin, average recording, quiet background, moderately variable, some disfluency, clear, wide pitch range, normal breath
(interest, awe, hope enthusiasm optimism · brisk, energised, neutral tension, casual) And you even have gradients going from the darker color to the lighter color at the bottom. And also they have really nice patterns. Now these are not just stock images, as when you click on them, you actually have a color ring at the top here, and when you click on that,
full caption & clip details
A child feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, slightly rough, slightly thin; clear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as interest, awe, hope enthusiasm optimism; style: casual, playful; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 1.7/10; 16.2s, EN.
EN_JPwc5UCV-bo_W000011 · in -19.7 dBFS · gain -0.3 dB · emolia-02229
(interest, contentment, concentration · normal-paced, very low-energy, slightly relaxed, casual) You can actually select what different colors are in the image. So, let's say we go to that first one. So, the triangles, like it changes the first row of triangles.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; clear, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as interest, contentment, concentration; style: casual, whispered; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.4/10; 10.6s, EN.
EN_JPwc5UCV-bo_W000012 · in -19.1 dBFS · gain -0.9 dB · emolia-02229
(hope enthusiasm optimism, elation, pleasure ecstasy · normal-paced, energised, neutral tension, casual) So let's say we want this to be like yellow, and then when you click on the other one, it actually changes the rest of the triangles for you. So I'm gonna have it at an orange color. So that's (breathy giggle) uh, that's really really nice. And also you can add text.
full caption & clip details
A child feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; clear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as hope enthusiasm optimism, elation, pleasure ecstasy; style: casual, playful; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 6.1/10; 13.2s, EN.
EN_JPwc5UCV-bo_W000013 · in -19.0 dBFS · gain -1.0 dB · emolia-02229
(pleasure ecstasy, sexual lust, embarrassment · normal-paced, energised, neutral tension, casual) Like that. So, I'm just gonna say this is a test. The animation is actually really nice. Now, I do have a few complaints. Like, there are lots of bugs. Do I wanna change the volume of clips?
full caption & clip details
A child feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; clear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as pleasure ecstasy, sexual lust, embarrassment; style: casual, conversational; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 4.8/10; 13.0s, EN.
EN_JPwc5UCV-bo_W000014 · in -19.0 dBFS · gain -1.0 dB · emolia-02229
(confusion, astonishment surprise, doubt · normal-paced, very low-energy, neutral tension, casual) It's actually such a hassle because when I turn it down to, let's say, a low volume, it's really laggy. You know, it wasn't laggy before, like this. Like, you see how laggy that is?
full caption & clip details
A child feminine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; clear, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as confusion, astonishment surprise, doubt; style: casual, playful; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 2.4/10; 10.0s, EN.
EN_JPwc5UCV-bo_W000015 · in -19.5 dBFS · gain -0.5 dB · emolia-02229
Concentration ↓  /  Confusionidentity −0.00 emotion 119 %   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #14

This chain comes from the proxy rule: the same two-sided test as above, but because Confusion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Confusion around average — 0.49, lower than 51 % of clips in this corpus — and ends with it strongly present at 0.81, higher than 81 % of clips in this corpus. That is a total rise of 0.32.

At the same time Concentration goes the other way, from 0.78 (higher than 78 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.24. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.10, then +0.23 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.89 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.89 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 30 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.751 before conversion and 0.746 after — it fell by 0.004. Neighbour-to-neighbour the worst pair went 0.770 → 0.764. (The earlier render, with segment 1 left raw, scores 0.806 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.324 in the original and +0.385 after conversion — 119 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.240 became -0.218.

Quality. Mean predicted overall quality across the segments went 2.92 → 3.07 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.751 → 0.746 -0.004identity cos neighbours 0.770 → 0.764d_b rescored +0.324 → +0.385d_a rescored -0.240 → -0.218d_a mined -0.239d_b mined 0.324min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_x0bcLTR3Pugtotal 29.5schain gain +2.5 dBseam step 0.3 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult somewhat masculine voice · neutral-bright, fairly smooth, balanced body, good recording, normal-paced, slightly relaxed, clear, light breath
(energised, moderately variable, some disfluency, casual) That is why it is called a global cooldown. It globally puts all skills on the cooldown, despite being different skills. All global cooldowns are
full caption & clip details
A young adult somewhat masculine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual; good recording, no background noise; mildly explicit content; genuineness 0.8/6; vocal-burst blend 0.1/10; 10.9s, EN.
EN_x0bcLTR3Pug_W000155 · in -20.0 dBFS · gain +0.0 dB · emolia-00808
(disgust, fear · energised, fairly steady, some disfluency, monologue) Off the global cooldown. They have longer timers and are usually marked as abilities. These can be used between GCDs. So in my case, True Thrust, Life Surge, Vorpal Thrust.
full caption & clip details
A child masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as disgust, fear; style: monologue, casual; good recording, quiet background; genuineness 0.3/6; vocal-burst blend 0.6/10; 13.8s, EN.
EN_x0bcLTR3Pug_W000156 · in -20.4 dBFS · gain +0.4 dB · emolia-00808
(normally alert, fairly steady, little disfluency, authoritative) And this has no pauses between attacks. No interruptions between GCDs.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, casual; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.9/10; 5.2s, EN.
EN_x0bcLTR3Pug_W000157 · in -18.9 dBFS · gain -1.1 dB · emolia-00808
Infatuation ↓  /  Sexual Lustidentity −0.03 emotion 117 %   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #15

This chain comes from the proxy rule: the same two-sided test as above, but because Sexual Lust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Sexual Lust around average — 0.46, lower than 54 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.43.

At the same time Infatuation goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.02, then +0.17, then +0.24 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 31 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.864 before conversion and 0.831 after — it fell by 0.033. Neighbour-to-neighbour the worst pair went 0.815 → 0.831. (The earlier render, with segment 1 left raw, scores 0.795 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sexual Lust moved +0.431 in the original and +0.505 after conversion — 117 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Infatuation, -0.280 became -0.072.

Quality. Mean predicted overall quality across the segments went 3.04 → 3.10 (+0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.864 → 0.831 -0.033identity cos neighbours 0.815 → 0.831d_b rescored +0.431 → +0.505d_a rescored -0.280 → -0.072d_a mined -0.307d_b mined 0.431min_cos_consec (site) 0.8851min_cos_anchor (site) 0.8851dataset emolialang zhspeaker ZH_B00052_S03080total 29.8schain gain -0.1 dBseam step 0.7 dBcrossfades 150/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a child feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, measured, normally alert
(infatuation · narration, storytelling) 这洁白无瑕的身子柔软又紧致,比皇宫中的美人还要好呢。
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as infatuation; style: narration, storytelling; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 2.5/10; 6.3s, ZH.
ZH_B00052_S03080_W000000 · in -12.7 dBFS · gain -7.3 dB · emolia-03792
(storytelling, whispered) 火燎闻言,猛然抬起头,微怒的眼睛看向柳依依。
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: storytelling, whispered; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 5.3/10; 4.2s, ZH.
ZH_B00052_S03080_W000001 · in -13.0 dBFS · gain -7.0 dB · emolia-03792
(narration, monologue) 而调皮的柳依依,一双娇好的大眼睛,微微眨动,一派天真无辜的样子。可是这个女人偏偏不脸红,依旧紧紧的盯着她的身子在看。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 2.8/10; 11.5s, ZH.
ZH_B00052_S03080_W000002 · in -13.8 dBFS · gain -6.2 dB · emolia-03792
(monologue, narration) 火燎俊美的容颜微微泛红,他越发的俊雅无邪,恰似雪山脚下开出的十里桃花,好不诱人。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 2.4/10; 8.4s, ZH.
ZH_B00052_S03080_W000003 · in -13.8 dBFS · gain -6.2 dB · emolia-03792
Contentment ↓  /  Prideidentity +0.07 emotion REVERSED   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #16

This chain comes from the proxy rule: the same two-sided test as above, but because Pride is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Pride around average — 0.53, higher than 53 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.32.

At the same time Contentment goes the other way, from 0.85 (higher than 85 % of clips in this corpus) to 0.63 (higher than 63 % of clips in this corpus), a change of -0.22. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then -0.07, then +0.14 — not a clean run: step 2 moves back the other way by 0.07 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.82 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.82 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 40 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.665 before conversion and 0.730 after — it rose by 0.065. Neighbour-to-neighbour the worst pair went 0.665 → 0.741. (The earlier render, with segment 1 left raw, scores 0.630 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.602 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Contentment, -0.222 became -0.639.

Quality. Mean predicted overall quality across the segments went 2.64 → 3.09 (+0.45) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.665 → 0.730 +0.065identity cos neighbours 0.665 → 0.741d_b rescored +0.602 → +0.000d_a rescored -0.222 → -0.639d_a mined -0.220d_b mined 0.315min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_ZtQxT3Zx9eEtotal 39.2schain gain +3.0 dBseam step 0.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, clear
(fairly steady, little disfluency, monologue, formal) Because Shannon is a former teacher, he has a unique perspective on the challenges that teachers face when integrating technology into the classroom.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.8/10; 7.0s, EN.
EN_ZtQxT3Zx9eE_W000029 · in -20.6 dBFS · gain +0.6 dB · emolia-01245
(concentration · steady, some disfluency, whispered, monologue) And that's why since the beginning, Shannon's felt that technology integration plays just as important of a role as the tech services in order for the consortium to be successful. So remember that seesaw example that I'd mentioned earlier?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: whispered, monologue; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 1.1/10; 12.3s, EN.
EN_ZtQxT3Zx9eE_W000030 · in -19.0 dBFS · gain -1.0 dB · emolia-01245
(steady, some disfluency, whispered, monologue) Before the consortium, there were no full-time integration specialists anywhere in Jackson County, but that's changed. Myself, Stacey Shue, and Brad Wilson now work together as a team of ed tech consultants who work directly with teachers in the consortium.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: whispered, monologue; average recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.1/10; 13.5s, EN.
EN_ZtQxT3Zx9eE_W000031 · in -20.9 dBFS · gain +0.9 dB · emolia-01245
(fairly steady, almost no disfluency, narration, monologue) Last year our team spent more than 2,000 face-to-face hours working with teachers and we expect that number to nearly double this year.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 1.0/10; 7.0s, EN.
EN_ZtQxT3Zx9eE_W000032 · in -18.9 dBFS · gain -1.1 dB · emolia-01245
Astonishment Surprise ↓  /  Intoxication Altered States of Consciousnessidentity −0.02 emotion 83 %   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #17

This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Intoxication Altered States of Consciousness clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.23.

At the same time Astonishment Surprise goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.20. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 7 s · snippets

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.662 before conversion and 0.646 after — it fell by 0.016. Neighbour-to-neighbour the worst pair went 0.662 → 0.646. (The earlier render, with segment 1 left raw, scores 0.543 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.231 in the original and +0.192 after conversion — 83 % of the delta retained, which is most of it. On the other named axis, Astonishment Surprise, -0.202 became -0.135.

Quality. Mean predicted overall quality across the segments went 2.57 → 2.77 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.662 → 0.646 -0.016identity cos neighbours 0.662 → 0.646d_b rescored +0.231 → +0.192d_a rescored -0.202 → -0.135d_a mined -0.201d_b mined 0.230min_cos_consec (site) —min_cos_anchor (site) —dataset snippetslang undspeaker batch260_part3_batch260_patotal 6.8schain gain +0.6 dBseam step 0.7 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normal-paced, normally alert, slightly relaxed
(conversational, playful) And what's the fish called? This one? Tunaika. Tunaika, okay.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational, playful; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 1.3/10; 3.8s.
batch260_part3_batch260_part3_chunk_80_1_471247 · in -17.1 dBFS · gain -2.9 dB · snippets-00837
(intoxication altered states of consciousness, sexual lust · casual, conversational) really expansive. And over here we have Sasha.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as intoxication altered states of consciousness, sexual lust; style: casual, conversational; good recording, no background noise; genuineness 3.4/6; vocal-burst blend 3.7/10; 3.1s.
batch260_part3_batch260_part3_chunk_80_1_471293 · in -15.9 dBFS · gain -4.1 dB · snippets-00837
Doubt ↓  /  Shameidentity −0.02 emotion 121 %   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #18

This chain comes from the proxy rule: the same two-sided test as above, but because Shame is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Shame clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.23.

At the same time Doubt goes the other way, from 0.93 (higher than 92 % of clips in this corpus) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.93 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.93 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 26 s · zh · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.904 before conversion and 0.882 after — it fell by 0.022. Neighbour-to-neighbour the worst pair went 0.904 → 0.882. (The earlier render, with segment 1 left raw, scores 0.854 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.233 in the original and +0.282 after conversion — 121 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Doubt, -0.249 became -0.162.

Quality. Mean predicted overall quality across the segments went 2.89 → 3.08 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.904 → 0.882 -0.022identity cos neighbours 0.904 → 0.882d_b rescored +0.233 → +0.282d_a rescored -0.249 → -0.162d_a mined -0.249d_b mined 0.233min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00007_S09869total 25.6schain gain +3.6 dBseam step 0.4 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, brisk, normally alert
(doubt · some disfluency, dramatic, didactic) 鹤岗官方呢已经回应了,房价呀并没有那么低。好的楼层呢,每平米呢两千到三千才是正常的价格,你想买个一百平的也得二十万左右。
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as doubt; style: dramatic, didactic; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 5.0/10; 9.7s, ZH.
ZH_B00007_S09869_W000001 · in -17.5 dBFS · gain -2.5 dB · emolia-03351
(almost no disfluency, monologue, didactic) 虽然呢比北上广深是便宜多了,但是呢也不是遍地白菜价。像小赵那种一两万的低价房子,一般呢是在楼梯房的顶楼呢或者非中心地段条件一般呢都不会太好。第二呀,鹤岗的气候环境呢,外地人呢不一定能适应。
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: monologue, didactic; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 5.8/10; 16.1s, ZH.
ZH_B00007_S09869_W000002 · in -17.9 dBFS · gain -2.1 dB · emolia-03351
Impatience and Irritability ↓  /  Fearidentity +0.09 emotion 34 %   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #19

This chain comes from the proxy rule: the same two-sided test as above, but because Fear is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fear around average — 0.49, right about the corpus median — and ends with it strongly present at 0.76, higher than 76 % of clips in this corpus. That is a total rise of 0.27.

At the same time Impatience and Irritability goes the other way, from 0.79 (higher than 79 % of clips in this corpus) to 0.56 (higher than 56 % of clips in this corpus), a change of -0.24. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.12, then +0.15 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.84 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.84 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 15 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.717 before conversion and 0.810 after — it rose by 0.092. Neighbour-to-neighbour the worst pair went 0.712 → 0.677. (The earlier render, with segment 1 left raw, scores 0.623 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.269 in the original and +0.091 after conversion — 34 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Impatience and Irritability, -0.238 became -0.176.

Quality. Mean predicted overall quality across the segments went 2.84 → 2.96 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.717 → 0.810 +0.092identity cos neighbours 0.712 → 0.677d_b rescored +0.269 → +0.091d_a rescored -0.238 → -0.176d_a mined -0.239d_b mined 0.269min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00006_S03888total 14.3schain gain +0.5 dBseam step 1.2 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(normal-paced, some disfluency, average clarity, authoritative) 我们恰恰就是因为这些年我们大规模开展了农村建设。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 3.5/10; 4.1s, ZH.
ZH_B00006_S03888_W000021 · in -18.5 dBFS · gain -1.5 dB · emolia-03340
(longing · measured, some disfluency, somewhat unclear, monologue) 比如说,二零零五年我们开展了新农村建设,到二零一七年又转型,为叫做什么叫做乡村振兴?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing; style: monologue, formal; average recording, no background noise; genuineness 2.6/6; vocal-burst blend 2.2/10; 6.4s, ZH.
ZH_B00006_S03888_W000022 · in -20.2 dBFS · gain +0.2 dB · emolia-03340
(normal-paced, no disfluency, clear, formal) 其实无外乎就是国家倾斜性的向乡村大规模投入。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 4.1/10; 4.2s, ZH.
ZH_B00006_S03888_W000023 · in -20.2 dBFS · gain +0.2 dB · emolia-03340
Distress ↓  /  Affectionidentity +0.07 emotion REVERSED   proxy_spearman__PXR__T0.20__C0.25__INTERNAL · #20

This chain comes from the proxy rule: the same two-sided test as above, but because Affection is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Affection clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it strongly present at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.23.

At the same time Distress goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.44 (lower than 56 % of clips in this corpus), a change of -0.52. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.94 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.94 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 12 s · zh · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.728 before conversion and 0.796 after — it rose by 0.068. Neighbour-to-neighbour the worst pair went 0.728 → 0.796. (The earlier render, with segment 1 left raw, scores 0.689 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Affection moved +0.581 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Distress, -0.519 became -0.517.

Quality. Mean predicted overall quality across the segments went 2.40 → 2.97 (+0.57) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.728 → 0.796 +0.068identity cos neighbours 0.728 → 0.796d_b rescored +0.581 → +0.000d_a rescored -0.519 → -0.517d_a mined -0.519d_b mined 0.229min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00010_S02724total 12.0schain gain +1.3 dBseam step 0.3 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · fairly smooth, thin, average recording, quiet background, normally alert, slightly relaxed, average clarity, moderate pitch range
(jealousy and envy, distress, contempt · measured, moderately variable, some disfluency, casual) 第四个误区就是追求老板喜欢,而不是用户喜欢。那么面对消费者的一个诉求呢,我们不要想着说。 (wistful sigh)
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is slightly cool, slightly dark, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as jealousy and envy, distress, contempt; style: casual, whispered; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 3.4/10; 7.6s, ZH.
ZH_B00010_S02724_W000022 · in -25.4 dBFS · gain +5.3 dB · emolia-03376
(fast, fairly steady, no disfluency, storytelling) 想当然他会喜欢了,我们一定要在他们各个平台上面的数据去判断。
full caption & clip details
A young adult feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: storytelling, dramatic; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 2.0/10; 4.6s, ZH.
ZH_B00010_S02724_W000023 · in -24.6 dBFS · gain +4.5 dB · emolia-03376