emotion_twosided__AB2__T0.25__C0.25__INTERNAL — voice-corrected

Manifest tier. emotion_twosided, rule AB2, T=0.25, step cap 0.25. Population 324,658 chains (3,589 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 260,655.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_emotion_twosided__AB2__T0.25__C0.25__INTERNAL.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
48segments re-voiced
0.830 → 0.819median worst-to-anchor identity cosine
74 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Concentration ↓  /  Prideidentity +0.05 emotion 24 %   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #1

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Pride up — by at least 0.25 each.

The chain starts with Pride around average — 0.53, higher than 53 % of clips in this corpus — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.36.

At the same time Concentration goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.17 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.79 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 20 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.778 before conversion and 0.830 after — it rose by 0.052. Neighbour-to-neighbour the worst pair went 0.758 → 0.779. (The earlier render, with segment 1 left raw, scores 0.504 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.356 in the original and +0.087 after conversion — 24 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Concentration, -0.299 became -0.254.

Quality. Mean predicted overall quality across the segments went 2.71 → 2.88 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.778 → 0.830 +0.052identity cos neighbours 0.758 → 0.779d_b rescored +0.356 → +0.087d_a rescored -0.299 → -0.254d_a mined -0.299d_b mined 0.356min_cos_consec (site) 0.7915min_cos_anchor (site) 0.7915dataset emolialang enspeaker EN_zx3jZ-R9oV8total 19.1schain gain +2.5 dBseam step 0.8 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(concentration, contemplation · monologue, didactic) But if I were to think in terms of sort of space, then this could equivalently represent the confirmation of a polymer.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, contemplation; style: monologue, didactic; good recording, quiet background; genuineness 3.7/6; vocal-burst blend 0.6/10; 6.3s, EN.
EN_zx3jZ-R9oV8_W000033 · in -21.7 dBFS · gain +1.7 dB · emolia-02585
(triumph, concentration · monologue, authoritative) What was time in my earlier case R square grew with time will be replaced by the number of monomers in your polymer in this case.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, concentration; style: monologue, authoritative; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 1.2/10; 8.6s, EN.
EN_zx3jZ-R9oV8_W000035 · in -18.1 dBFS · gain -1.9 dB · emolia-02585
(casual, monologue) So that is what we will start off with. With the most basic model in some sense of a polymer
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 0.8/10; 4.6s, EN.
EN_zx3jZ-R9oV8_W000036 · in -19.9 dBFS · gain -0.1 dB · emolia-02585
Disgust ↓  /  Impatience and Irritabilityidentity −0.01 emotion 63 %   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #2

This chain comes from the two-sided rule: it only counts if both emotions move — Disgust down and Impatience and Irritability up — by at least 0.25 each.

The chain starts with Impatience and Irritability clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.31.

At the same time Disgust goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.06, then +0.25 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 39 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.784 before conversion and 0.774 after — it fell by 0.011. Neighbour-to-neighbour the worst pair went 0.784 → 0.774. (The earlier render, with segment 1 left raw, scores 0.652 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.310 in the original and +0.195 after conversion — 63 % of the delta retained. On the other named axis, Disgust, -0.405 became -0.185.

Quality. Mean predicted overall quality across the segments went 2.97 → 3.11 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.784 → 0.774 -0.011identity cos neighbours 0.784 → 0.774d_b rescored +0.310 → +0.195d_a rescored -0.405 → -0.185d_a mined -0.404d_b mined 0.310min_cos_consec (site) 0.8862min_cos_anchor (site) 0.8862dataset emolialang enspeaker EN_4UVPsQIAbTctotal 38.3schain gain +5.3 dBseam step 1.9 dBcrossfades 150/100 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, good recording, some disfluency, average clarity, moderate pitch range, light breath
(disgust · normal-paced, normally alert, slightly relaxed, casual) (low mumble) They put all these African American mayors up there, (low mumble) um, instead of the activists who were yelling at the mayors.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as disgust; style: casual, conversational; good recording, quiet background; genuineness 3.5/6; vocal-burst blend 1.8/10; 7.0s, EN.
EN_4UVPsQIAbTc_W000198 · in -20.4 dBFS · gain +0.4 dB · emolia-00367
(sourness, contempt, concentration · brisk, normally alert, slightly relaxed, casual) (ahem) Um, and I was thinking, well, yeah, of course. Like, you put up, (ahem) uh, responsible, accountable, elected officials who, (low mumble) um, you know, have to manage the various considerations that are up there, not activists yelling in the streets. And I feel like Seattle and Portland don't have the kind of political institutions and political stakeholders who can say no.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as sourness, contempt, concentration; style: casual, conversational; good recording, quiet background; genuineness 3.1/6; vocal-burst blend 5.0/10; 21.8s, EN.
EN_4UVPsQIAbTc_W000199 · in -21.6 dBFS · gain +1.6 dB · emolia-00367
(impatience and irritability, fear, disappointment · brisk, energised, neutral tension, casual) Like, this is not helping. Like, I agree that there's a problem here and I want to work on it. But the specific thing that you guys are doing right now in the streets is not in any way.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as impatience and irritability, fear, disappointment; style: casual, dramatic; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 4.1/10; 10.0s, EN.
EN_4UVPsQIAbTc_W000200 · in -23.3 dBFS · gain +3.3 dB · emolia-00367
Pride ↓  /  Angeridentity +0.01 emotion 78 %   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #3

This chain comes from the two-sided rule: it only counts if both emotions move — Pride down and Anger up — by at least 0.25 each.

The chain starts with Anger around average — 0.58, higher than 58 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.37.

At the same time Pride goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.25, then -0.07, then -0.00, then +0.19 — not a clean run: step 2 moves back the other way by 0.07 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 72 s · bg · eurospeech

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.914 before conversion and 0.923 after — it rose by 0.009. Neighbour-to-neighbour the worst pair went 0.931 → 0.934. (The earlier render, with segment 1 left raw, scores 0.669 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Anger moved +0.369 in the original and +0.288 after conversion — 78 % of the delta retained, which is most of it. On the other named axis, Pride, -0.275 became -0.166.

Quality. Mean predicted overall quality across the segments went 3.21 → 3.31 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.914 → 0.923 +0.009identity cos neighbours 0.931 → 0.934d_b rescored +0.369 → +0.288d_a rescored -0.275 → -0.166d_a mined -0.275d_b mined 0.369min_cos_consec (site) 0.9343min_cos_anchor (site) 0.9301dataset eurospeechlang bgspeaker bulgaria_bulgaria_1_070220total 70.5schain gain +0.7 dBseam step 0.5 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a child masculine voice · thin, average recording, neutral tension, moderately variable
(pride, disgust · fast, energised, some disfluency, storytelling) по кардиостимулация и електрофизиология и група пациенти, които от години се борят за цялостната реимбурсация на имплантируем кардиовертер дефибрилатор от
full caption & clip details
A child masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, dark, slightly rough, thin; slurred, some disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as pride, disgust; style: storytelling, cartoonish; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 6.6/10; 13.2s, BG.
bulgaria_bulgaria_1_07022019_3550752_3563904 · in -27.6 dBFS · gain +7.7 dB · eurospeech-00087
(bitterness, affection · measured, energised, frequent disfluency, monologue) но до момента не са срещнали разбиране от институциите и няма резултат по отношение на пълната реимбурсация или подобряване на съществуващата, за да се намали до минимум доплащането от страна на пациентите.
full caption & clip details
An elderly masculine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is slightly cool, dark, rough, thin; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as bitterness, affection; style: monologue, storytelling; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 7.1/10; 14.6s, BG.
bulgaria_bulgaria_1_07022019_3563904_3578480 · in -27.4 dBFS · gain +7.3 dB · eurospeech-00087
(malevolence malice · normal-paced, normally alert, almost no disfluency, monologue) В момента Националната здравноосигурителна каса заплаща на 50% от устройството, а другите 50% – пациентът, като сумата е между 3000 и 5000 лв.,
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as malevolence malice; style: monologue, storytelling; average recording, no background noise; genuineness 1.5/6; vocal-burst blend 3.8/10; 10.9s, BG.
bulgaria_bulgaria_1_07022019_3578480_3589344 · in -27.7 dBFS · gain +7.7 dB · eurospeech-00087
(thankfulness gratitude, disgust, malevolence malice · measured, energised, some disfluency, storytelling) което е непосилно за повечето хора. Имплантируемият кардиовертер дефибрилатор е най сигурният и доказан начин за профилактика на внезапната сърдечна смърт в целия свят.
full caption & clip details
An elderly masculine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is slightly cool, slightly dark, slightly rough, thin; somewhat unclear, some disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as thankfulness gratitude, disgust, malevolence malice; style: storytelling, cartoonish; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 5.1/10; 13.6s, BG.
bulgaria_bulgaria_1_07022019_3589344_3602975 · in -28.3 dBFS · gain +8.3 dB · eurospeech-00087
(anger, disappointment · normal-paced, normally alert, some disfluency, monologue) За съжаление, страната ни е на едно от последните места по имплантиране на такъв дефибрилатор. Въпросът ми към Вас е: ще бъде ли отделена необходимата сума от бюджета на Националната здравноосигурителна каса за пълната реимбурсация на имплантируемия кардиовертер дефибрилатор?
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, rough, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger, disappointment; style: monologue, cartoonish; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 5.8/10; 19.2s, BG.
bulgaria_bulgaria_1_07022019_3602975_3622127 · in -27.9 dBFS · gain +7.9 dB · eurospeech-00087
Doubt ↓  /  Contentmentidentity +0.89 emotion 114 %   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #4

This chain comes from the two-sided rule: it only counts if both emotions move — Doubt down and Contentment up — by at least 0.25 each.

The chain starts with Contentment clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.28.

At the same time Doubt goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.10, then -0.04, then +0.22 — not a clean run: step 2 moves back the other way by 0.04 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.05 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.09 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.05, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 74 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.063 before conversion and 0.827 after — it rose by 0.891. Neighbour-to-neighbour the worst pair went -0.149 → 0.725. (The earlier render, with segment 1 left raw, scores 0.567 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.277 in the original and +0.315 after conversion — 114 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Doubt, -0.307 became -0.081.

Quality. Mean predicted overall quality across the segments went 2.58 → 3.21 (+0.62) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 -0.063 → 0.827 +0.891identity cos neighbours -0.149 → 0.725d_b rescored +0.277 → +0.315d_a rescored -0.307 → -0.081d_a mined -0.309d_b mined 0.278min_cos_consec (site) -0.0912min_cos_anchor (site) -0.0514dataset podcastlang enspeaker 115163total 73.3schain gain +5.1 dBseam step 1.4 dBcrossfades 150/150/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, slightly rough, quiet background, measured, somewhat unclear, fairly narrow pitch
(doubt, confusion, concentration · very low-energy, relaxed, fairly steady, casual) let's say someone listening to this, whether it's a within our church or maybe it's just somebody in the area, (low mumble) um wants to get involved with this kind of this type of ministry. Where's a good entry point for them? Is it like come talk to Danny or is there like a website they could get information on, or is there (low mumble) um something
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, slightly guarded; reads as doubt, confusion, concentration; style: casual, monologue; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 7.0/10; 22.4s, EN.
115163_00304604 · in -33.2 dBFS · gain +13.2 dB · podcast-02648
(doubt, fear · very low-energy, relaxed, fairly steady, casual) that they would need to do? (low mumble) Um so if somebody's interested in maybe that (low mumble) um helping (low mumble) um, you know with the transitional with a someone paroling out, getting them a (low mumble) you know, helping them find a job or helping I mean maybe you know, there might be someone that's interested in that. Um
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, neutral stance, slightly guarded; reads as doubt, fear; style: casual, monologue; below-average recording, quiet background; genuineness 5.4/6; vocal-burst blend 8.9/10; 19.0s, EN.
115163_00306840 · in -34.9 dBFS · gain +14.8 dB · podcast-02650
(very low-energy, relaxed, fairly steady, casual) demand it so much. (low mumble) Um okay. So I guess if you're if you are a listener that is within MCO, talk to Danny. And if you're a listener that is not, maybe just call the church. Call
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, whispered; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 5.9/10; 12.9s, EN.
115163_00310856 · in -31.0 dBFS · gain +11.0 dB · podcast-02838
(contentment, thankfulness gratitude, triumph · subdued, slightly relaxed, steady, casual) place to plug in because what we're doing is we are getting the guys that are a part of celebrate recovery inside. So that when they come out, they are already established somewhere and have someplace to go.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as contentment, thankfulness gratitude, triumph; style: casual, monologue; below-average recording, quiet background; genuineness 3.5/6; vocal-burst blend 4.0/10; 19.6s, EN.
115163_00312632 · in -25.8 dBFS · gain +5.8 dB · podcast-02841
Concentration ↓  /  Astonishment Surpriseidentity −0.01 emotion REVERSED   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #5

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Astonishment Surprise up — by at least 0.25 each.

The chain starts with Astonishment Surprise below average — 0.36, lower than 64 % of clips in this corpus — and ends with it clearly present at 0.66, higher than 66 % of clips in this corpus. That is a total rise of 0.30.

At the same time Concentration goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.63 (higher than 63 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.13 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 24 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.823 before conversion and 0.816 after — it fell by 0.006. Neighbour-to-neighbour the worst pair went 0.848 → 0.818. (The earlier render, with segment 1 left raw, scores 0.685 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.300 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Concentration, -0.321 became -0.365.

Quality. Mean predicted overall quality across the segments went 2.77 → 3.01 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.823 → 0.816 -0.006identity cos neighbours 0.848 → 0.818d_b rescored +0.300 → +0.000d_a rescored -0.321 → -0.365d_a mined -0.321d_b mined 0.300min_cos_consec (site) 0.8343min_cos_anchor (site) 0.8499dataset emolialang enspeaker EN_tD_P9HCODS4total 23.4schain gain +1.5 dBseam step 1.4 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, fairly steady, average clarity, moderate pitch range
(concentration · normal-paced, neutral tension, some disfluency, casual) Do add one to the plain buttons counter. So you're just gonna go ahead and do that. That'll be one event. From there, you're gonna say, when, (ahem) (low mumble) uh, the button counter is one,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 4.5/10; 10.4s, EN.
EN_tD_P9HCODS4_W000002 · in -20.4 dBFS · gain +0.5 dB · emolia-01151
(measured, slightly relaxed, frequent disfluency, monologue) Do the camera's position should go to camera point one, scene one. Do when, (low mumble) uh, it's counter is two, go to scene two.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.5/10; 8.6s, EN.
EN_tD_P9HCODS4_W000003 · in -19.0 dBFS · gain -1.0 dB · emolia-01151
(brisk, slightly relaxed, some disfluency, casual) And then after 2, there's no more scene, so just set a condition that says, hey, when the counter is greater than 2,
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 1.4/10; 4.7s, EN.
EN_tD_P9HCODS4_W000004 · in -21.9 dBFS · gain +1.9 dB · emolia-01151
Shame ↓  /  Fatigue Exhaustionidentity −0.01 emotion 104 %   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #6

This chain comes from the two-sided rule: it only counts if both emotions move — Shame down and Fatigue Exhaustion up — by at least 0.25 each.

The chain starts with Fatigue Exhaustion around average — 0.47, lower than 53 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.38.

At the same time Shame goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.35. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 28 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.907 before conversion and 0.900 after — it fell by 0.007. Neighbour-to-neighbour the worst pair went 0.866 → 0.862. (The earlier render, with segment 1 left raw, scores 0.829 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.376 in the original and +0.392 after conversion — 104 % of the delta retained, which is essentially all of it. On the other named axis, Shame, -0.345 became -0.327.

Quality. Mean predicted overall quality across the segments went 3.03 → 3.24 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.907 → 0.900 -0.007identity cos neighbours 0.866 → 0.862d_b rescored +0.376 → +0.392d_a rescored -0.345 → -0.327d_a mined -0.345d_b mined 0.376min_cos_consec (site) 0.9034min_cos_anchor (site) 0.9164dataset emolialang zhspeaker ZH_B00013_S08651total 27.0schain gain +3.1 dBseam step 2.2 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, normally alert, slightly relaxed, fairly steady
(shame, confusion, embarrassment · fast, average clarity, didactic, monologue) 就因为它本身其实并没有什么难度,并不需要真的是让你把上百个这种拼图的碎片,然后从里面找出最有用的几个,其实不是这个意思。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, confusion, embarrassment; style: didactic, monologue; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 6.5/10; 10.2s, ZH.
ZH_B00013_S08651_W000039 · in -27.5 dBFS · gain +7.5 dB · emolia-03403
(astonishment surprise, relief · measured, somewhat unclear, monologue, didactic) 哎,但是这个设计我觉得很好,就是他辅助了帮助了我这样的人,不怎么习惯去动脑子,去琢磨游戏的人,让我可以很好的去思考。然后最后呢相对来说啊比较。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as astonishment surprise, relief; style: monologue, didactic; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 6.3/10; 12.0s, ZH.
ZH_B00013_S08651_W000040 · in -29.9 dBFS · gain +9.9 dB · emolia-03403
(fast, average clarity, casual, conversational) (tsk) (ahem) 呃,清晰的去得出这样一个整个的一个结论,啊,把这个案情了解的更透彻一些。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: casual, conversational; average recording, no background noise; genuineness 4.8/6; vocal-burst blend 8.0/10; 5.2s, ZH.
ZH_B00013_S08651_W000041 · in -30.2 dBFS · gain +10.2 dB · emolia-03403
Contemplation ↓  /  Fatigue Exhaustionidentity −0.03 emotion REVERSED   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #7

This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Fatigue Exhaustion up — by at least 0.25 each.

The chain starts with Fatigue Exhaustion around average — 0.52, higher than 52 % of clips in this corpus — and ends with it strongly present at 0.81, higher than 81 % of clips in this corpus. That is a total rise of 0.29.

At the same time Contemplation goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.35. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.10 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 28 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.720 before conversion and 0.686 after — it fell by 0.034. Neighbour-to-neighbour the worst pair went 0.720 → 0.686. (The earlier render, with segment 1 left raw, scores 0.540 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Fatigue Exhaustion moved +0.293 in the original and -0.676 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Contemplation, -0.349 became -0.262.

Quality. Mean predicted overall quality across the segments went 2.55 → 2.83 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.720 → 0.686 -0.034identity cos neighbours 0.720 → 0.686d_b rescored +0.293 → -0.676d_a rescored -0.349 → -0.262d_a mined -0.349d_b mined 0.293min_cos_consec (site) 0.8832min_cos_anchor (site) 0.8892dataset emolialang enspeaker EN_MJetgOvgrpItotal 27.6schain gain -2.2 dBseam step 5.0 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an elderly masculine voice · neutral-toned, fairly smooth, average recording, quiet background, fairly steady
(contemplation, concentration · slow, very low-energy, relaxed, ASMR) (low mumble) Uhm, how were you satisfied that it was ready to go live? (ahem) Uhm, and we've talked about governance, (low mumble) uhm, you know, an organization's readiness to accept responsibility for how its machine learning systems behave and to be able to respond, you know, in a timely way to
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, neutral openness; reads as contemplation, concentration; style: ASMR, monologue; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 5.1/10; 18.5s, EN.
EN_MJetgOvgrpI_W000192 · in -18.0 dBFS · gain -2.0 dB · emolia-02599
(emotional numbness · normal-paced, very low-energy, slightly relaxed, whispered) Subjects, you know, citizens or customers who raise concerns or raise a complaint.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, slightly thin; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as emotional numbness; style: whispered, casual; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.8/10; 4.7s, EN.
EN_MJetgOvgrpI_W000193 · in -14.6 dBFS · gain -5.4 dB · emolia-02599
(measured, subdued, slightly relaxed, monologue) (ahem) Uhm, and being accessible, you know, to many organizations, that isn't accessible, so.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 1.4/10; 4.9s, EN.
EN_MJetgOvgrpI_W000194 · in -14.7 dBFS · gain -5.3 dB · emolia-02599
Fear ↓  /  Fatigue Exhaustionidentity +0.05 emotion 77 %   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #8

This chain comes from the two-sided rule: it only counts if both emotions move — Fear down and Fatigue Exhaustion up — by at least 0.25 each.

The chain starts with Fatigue Exhaustion around average — 0.42, lower than 58 % of clips in this corpus — and ends with it clearly present at 0.75, higher than 75 % of clips in this corpus. That is a total rise of 0.33.

At the same time Fear goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.49 (right about the corpus median), a change of -0.46. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.10, then +0.23 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 23 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.858 before conversion and 0.907 after — it rose by 0.048. Neighbour-to-neighbour the worst pair went 0.901 → 0.897. (The earlier render, with segment 1 left raw, scores 0.864 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.350 in the original and +0.268 after conversion — 77 % of the delta retained, which is most of it. On the other named axis, Fear, -0.480 became -0.099.

Quality. Mean predicted overall quality across the segments went 3.04 → 3.21 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.858 → 0.907 +0.048identity cos neighbours 0.901 → 0.897d_b rescored +0.350 → +0.268d_a rescored -0.480 → -0.099d_a mined -0.457d_b mined 0.328min_cos_consec (site) 0.9265min_cos_anchor (site) 0.8920dataset emolialang zhspeaker ZH_B00005_S08484total 22.3schain gain +2.9 dBseam step 0.5 dBcrossfades 100/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady
(fear · measured, narration, authoritative) 因为盖长公主主张这么办,他也不便于过于固执。可是封丁外人为侯,算是什么规矩呢?霍光就是不易。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear; style: narration, authoritative; average recording, no background noise; genuineness 1.6/6; vocal-burst blend 5.8/10; 9.2s, ZH.
ZH_B00005_S08484_W000015 · in -19.3 dBFS · gain -0.7 dB · emolia-03326
(relief · fast, authoritative, narration) 上官桀,他们勾结燕王汉昭帝的义母哥哥刘旦,先想办法消灭霍光,然后废去汉昭帝立燕王刘旦为皇帝。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief; style: authoritative, narration; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 4.9/10; 8.7s, ZH.
ZH_B00005_S08484_W000016 · in -16.6 dBFS · gain -3.4 dB · emolia-03326
(normal-paced, authoritative, formal) 朝廷里有左将军、上官杰车骑将军上官安,还有别的大臣。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.5/10; 4.8s, ZH.
ZH_B00005_S08484_W000017 · in -15.3 dBFS · gain -4.7 dB · emolia-03326
Fatigue Exhaustion ↓  /  Affectionidentity +0.29 emotion 52 %   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #9

This chain comes from the two-sided rule: it only counts if both emotions move — Fatigue Exhaustion down and Affection up — by at least 0.25 each.

The chain starts with Affection clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.40.

At the same time Fatigue Exhaustion goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.18, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.26 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.28 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.26, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 31 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.203 before conversion and 0.496 after — it rose by 0.293. Neighbour-to-neighbour the worst pair went 0.266 → 0.588. (The earlier render, with segment 1 left raw, scores 0.350 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Affection moved +0.404 in the original and +0.210 after conversion — 52 % of the delta retained. On the other named axis, Fatigue Exhaustion, -0.290 became -0.436.

Quality. Mean predicted overall quality across the segments went 2.87 → 3.20 (+0.33) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.203 → 0.496 +0.293identity cos neighbours 0.266 → 0.588d_b rescored +0.404 → +0.210d_a rescored -0.290 → -0.436d_a mined -0.291d_b mined 0.404min_cos_consec (site) 0.2765min_cos_anchor (site) 0.2554dataset podcastlang enspeaker 449725total 30.6schain gain +3.7 dBseam step 2.1 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, balanced body, average recording, moderately variable, frequent disfluency
(fatigue exhaustion, triumph, elation · normal-paced, normally alert, neutral tension, casual) I'm not mad about it. (low mumble) Um we got one last topic we want to hit tonight. We had two, actually. We've got a coming soon album that you need to check out when it comes out. The Danny Brown album.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as fatigue exhaustion, triumph, elation; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 4.8/10; 12.3s, EN.
449725_00110263 · in -27.9 dBFS · gain +7.9 dB · podcast-05468
(pleasure ecstasy, interest, intoxication altered states of consciousness · normal-paced, very low-energy, neutral tension, casual) Oh, that's right. The Danny Brown album. (ahem) Uh if you don't know Danny Brown, I recommend listening to the album XXX. Came out in two thousand twelve. Really cool album. You could say it's shock rap. It's it's different. It's it's not like you're
full caption & clip details
An adult masculine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as pleasure ecstasy, interest, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.3/6; vocal-burst blend 3.1/10; 14.2s, EN.
449725_00111519 · in -31.4 dBFS · gain +11.4 dB · podcast-05448
(affection, relief, thankfulness gratitude · slow, very low-energy, relaxed, casual) God bless you. God bless all of you. Anyway, um
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is neutral-toned, dark, slightly rough, balanced body; slurred, frequent disfluency, wide pitch range, audible breath; affect is mildly negative, submissive, neutral openness; reads as affection, relief, thankfulness gratitude; style: casual, whispered; average recording, no background noise; genuineness 3.8/6; vocal-burst blend 2.2/10; 4.4s, EN.
449725_00113591 · in -32.6 dBFS · gain +12.6 dB · podcast-05448
Disgust ↓  /  Prideidentity −0.02 emotion REVERSED   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #10

This chain comes from the two-sided rule: it only counts if both emotions move — Disgust down and Pride up — by at least 0.25 each.

The chain starts with Pride clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.26.

At the same time Disgust goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.11, then +0.14 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 37 s · pt · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.844 before conversion and 0.821 after — it fell by 0.024. Neighbour-to-neighbour the worst pair went 0.785 → 0.768. (The earlier render, with segment 1 left raw, scores 0.734 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.255 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Disgust, -0.329 became -0.648.

Quality. Mean predicted overall quality across the segments went 2.89 → 3.27 (+0.38) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.844 → 0.821 -0.024identity cos neighbours 0.785 → 0.768d_b rescored +0.255 → +0.000d_a rescored -0.329 → -0.648d_a mined -0.329d_b mined 0.256min_cos_consec (site) 0.8580min_cos_anchor (site) 0.8641dataset podcastlang ptspeaker 899254total 36.4schain gain -0.3 dBseam step 5.1 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, average recording, normally alert
(disgust, shame, contempt · measured, neutral tension, fairly steady, casual) É quase tão gente como a moda dos americanos que eu já tinha falado outra vez que é B leite com aquele puré de batata e aquele bitoque e aquele ovo estrelado. É pá, não.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, thin; somewhat unclear, some disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, shame, contempt; style: casual, monologue; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 7.1/10; 10.2s, PT.
899254_00082672 · in -21.1 dBFS · gain +1.1 dB · podcast-03633
(elation, embarrassment, infatuation · normal-paced, neutral tension, moderately variable, casual) É uma cena que nem vos passa pela cabeça do tipo, ah, tem que ser reais, não há leito, que é gay, não há iogurto, que é gay. O quê? Tem água da torneira? Pois tenho, mas não vou usá-la, de certeza, vai para o caralho. Estou meado, vou mesmo meado. Cereisinhos, está, tal, abrir água da torneira, até podes meter tipo quentinho. É quase tão grave como
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as elation, embarrassment, infatuation; style: casual, cartoonish; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 7.4/10; 20.9s, PT.
899254_00083684 · in -18.4 dBFS · gain -1.6 dB · podcast-05922
(normal-paced, slightly relaxed, fairly steady, casual) (ahem) isto é, o nível a seguir de aquecer o leite com os cereais.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, no background noise; genuineness 4.5/6; vocal-burst blend 3.8/10; 5.7s, PT.
899254_00085776 · in -20.4 dBFS · gain +0.4 dB · podcast-05943
Shame ↓  /  Affectionidentity +0.02 emotion 42 %   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #11

This chain comes from the two-sided rule: it only counts if both emotions move — Shame down and Affection up — by at least 0.25 each.

The chain starts with Affection clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.35.

At the same time Shame goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.19 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 57 s · de · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.911 before conversion and 0.932 after — it rose by 0.021. Neighbour-to-neighbour the worst pair went 0.908 → 0.922. (The earlier render, with segment 1 left raw, scores 0.753 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Affection moved +0.351 in the original and +0.147 after conversion — 42 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Shame, -0.253 became -0.055.

Quality. Mean predicted overall quality across the segments went 2.96 → 3.15 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.911 → 0.932 +0.021identity cos neighbours 0.908 → 0.922d_b rescored +0.351 → +0.147d_a rescored -0.253 → -0.055d_a mined -0.253d_b mined 0.351min_cos_consec (site) 0.9273min_cos_anchor (site) 0.9238dataset emolialang despeaker DE_DUihznsHYIytotal 56.1schain gain +3.0 dBseam step 0.6 dBcrossfades 100/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, normal-paced, normally alert, fairly steady
(shame, sadness, concentration · slightly relaxed, casual, monologue) Mein Eindruck ist, und der Eindruck, glaube ich, den Eindruck haben viele, dass je größer die Konkurrenz auf dem Wohnungsmarkt ist und je mehr, (low mumble) ähm, oder sozusagen, desto schwieriger ist es ja für Menschen mit niedrigem Einkommen, für Menschen, die Diskriminierungserfahrung auch haben.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as shame, sadness, concentration; style: casual, monologue; average recording, no background noise; genuineness 4.0/6; vocal-burst blend 0.0/10; 16.3s, DE.
DE_DUihznsHYIy_W000284 · in -22.1 dBFS · gain +2.1 dB · emolia-00186
(fear, disappointment, helplessness · slightly relaxed, monologue, casual) (ahem) ähm, eben an eine Wohnung zu kommen und dass sich da sozusagen Fragen von, (ahem) äh, des, sozusagen unterschiedlichen Diskriminierungserfahrungen und der Frage auch von sozialer, sozusagen, Benachteiligung oder Marginalisierung dann eben gerade auf dem Wohnungsmarkt auch nochmal, sagen wir, zusammenkommen, doppeln, (ahem) ähm, oder verdreifachen und das deshalb umso schwieriger wird, eine Wohnung, eine Wohnung zu kommen. Deshalb ist, ist das glaube ich eine ganz wichtige Aufgabe,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as fear, disappointment, helplessness; style: monologue, casual; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 0.0/10; 25.3s, DE.
DE_DUihznsHYIy_W000285 · in -19.9 dBFS · gain -0.1 dB · emolia-00186
(affection, relief · neutral tension, casual, monologue) Ich mache mal hier Schluss, damit die anderen auch noch reden können.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as affection, relief; style: casual, monologue; average recording, no background noise; genuineness 4.8/6; vocal-burst blend 1.5/10; 14.9s, DE.
DE_DUihznsHYIy_W000286 · in -20.9 dBFS · gain +0.9 dB · emolia-00186
Impatience and Irritability ↓  /  Astonishment Surpriseidentity −0.03 emotion 145 %   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #12

This chain comes from the two-sided rule: it only counts if both emotions move — Impatience and Irritability down and Astonishment Surprise up — by at least 0.25 each.

The chain starts with Astonishment Surprise clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.30.

At the same time Impatience and Irritability goes the other way, from 0.83 (higher than 83 % of clips in this corpus) to 0.48 (lower than 52 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.18 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 32 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.787 before conversion and 0.755 after — it fell by 0.032. Neighbour-to-neighbour the worst pair went 0.787 → 0.755. (The earlier render, with segment 1 left raw, scores 0.685 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.304 in the original and +0.441 after conversion — 145 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Impatience and Irritability, -0.356 became -0.348.

Quality. Mean predicted overall quality across the segments went 2.88 → 3.14 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.787 → 0.755 -0.032identity cos neighbours 0.787 → 0.755d_b rescored +0.304 → +0.441d_a rescored -0.356 → -0.348d_a mined -0.356d_b mined 0.304min_cos_consec (site) 0.8310min_cos_anchor (site) 0.8305dataset emolialang enspeaker EN_B00028_S05392total 31.3schain gain +2.4 dBseam step 0.5 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged feminine voice · neutral-bright, good recording, no background noise, slightly relaxed, clear
(slow, normally alert, steady, narration) You mustn't make any sudden moves or make a sound. And above all, you mustn't run.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, little disfluency, moderate pitch range, audible breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: narration, whispered; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.3/10; 8.1s, EN.
EN_B00028_S05392_W000015 · in -20.8 dBFS · gain +0.8 dB · emolia-00785
(fear, pain, helplessness · measured, subdued, fairly steady, whispered) No one can run faster in the forest than a bear. And remember, we don't have a gun to keep us safe. That night we went to sleep. Or we tried to.
full caption & clip details
A child feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; reads as fear, pain, helplessness; style: whispered, didactic; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.1/10; 13.5s, EN.
EN_B00028_S05392_W000016 · in -20.6 dBFS · gain +0.6 dB · emolia-00785
(astonishment surprise · normal-paced, normally alert, fairly steady, whispered) The next day we stopped at eleven o'clock for a break. While the others were resting I went for a walk in the forest.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as astonishment surprise; style: whispered, narration; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.3/10; 10.1s, EN.
EN_B00028_S05392_W000017 · in -19.4 dBFS · gain -0.6 dB · emolia-00785
Astonishment Surprise ↓  /  Intoxication Altered States of Consciousnessidentity −0.04 emotion 87 %   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #13

This chain comes from the two-sided rule: it only counts if both emotions move — Astonishment Surprise down and Intoxication Altered States of Consciousness up — by at least 0.25 each.

The chain starts with Intoxication Altered States of Consciousness below average — 0.40, lower than 60 % of clips in this corpus — and ends with it clearly present at 0.66, higher than 66 % of clips in this corpus. That is a total rise of 0.26.

At the same time Astonishment Surprise goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.13 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 34 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.653 before conversion and 0.611 after — it fell by 0.042. Neighbour-to-neighbour the worst pair went 0.653 → 0.611. (The earlier render, with segment 1 left raw, scores 0.550 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.265 in the original and +0.232 after conversion — 87 % of the delta retained, which is most of it. On the other named axis, Astonishment Surprise, -0.329 became -0.334.

Quality. Mean predicted overall quality across the segments went 2.79 → 2.89 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.653 → 0.611 -0.042identity cos neighbours 0.653 → 0.611d_b rescored +0.265 → +0.232d_a rescored -0.329 → -0.334d_a mined -0.329d_b mined 0.265min_cos_consec (site) 0.8475min_cos_anchor (site) 0.8318dataset emolialang enspeaker EN_B00054_S03669total 33.5schain gain +1.1 dBseam step 1.7 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, good recording, slightly relaxed, moderate pitch range, light breath
(astonishment surprise, teasing · normal-paced, normally alert, fairly steady, storytelling) A male realizes that he's totally gonna get the chance to mate.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as astonishment surprise, teasing; style: storytelling, casual; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 2.7/10; 3.0s, EN.
EN_B00054_S03669_W000033 · in -23.7 dBFS · gain +3.7 dB · emolia-01293
(sourness, disgust, interest · brisk, energised, moderately variable, casual) At this point, spongy tissue in the penis fills with blood, and bam, erection. Some animals like raccoons, whales, and walruses actually have a literal bone in their penis to help the erection along. But either way, the point is to allow the penis to enter the vagina, which scientists call coitus, and deposit the sperm he's put so much into making. These sperm travel in a special fluid, semen, whose ingredients aren't combined until they're ready to be released by a series of muscular contractions that cause emission, more commonly known as ejaculation.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as sourness, disgust, interest; style: casual, dramatic; good recording, quiet background; genuineness 0.4/6; vocal-burst blend 3.4/10; 27.4s, EN.
EN_B00054_S03669_W000034 · in -20.2 dBFS · gain +0.2 dB · emolia-01293
(brisk, normally alert, fairly steady, casual) At this point, the contractions carry the mature sperm from the epididymis
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 1.8/10; 3.5s, EN.
EN_B00054_S03669_W000035 · in -21.1 dBFS · gain +1.1 dB · emolia-01293
Fatigue Exhaustion ↓  /  Malevolence Maliceidentity −0.11 emotion 94 %   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #14

This chain comes from the two-sided rule: it only counts if both emotions move — Fatigue Exhaustion down and Malevolence Malice up — by at least 0.25 each.

The chain starts with Malevolence Malice below average — 0.35, lower than 65 % of clips in this corpus — and ends with it clearly present at 0.70, higher than 70 % of clips in this corpus. That is a total rise of 0.35.

At the same time Fatigue Exhaustion goes the other way, from 0.83 (higher than 83 % of clips in this corpus) to 0.57 (higher than 57 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.13 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 18 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.917 before conversion and 0.805 after — it fell by 0.111. Neighbour-to-neighbour the worst pair went 0.939 → 0.805. (The earlier render, with segment 1 left raw, scores 0.620 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Malevolence Malice moved +0.354 in the original and +0.332 after conversion — 94 % of the delta retained, which is essentially all of it. On the other named axis, Fatigue Exhaustion, -0.256 became -0.235.

Quality. Mean predicted overall quality across the segments went 2.84 → 2.89 (+0.05) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.917 → 0.805 -0.111identity cos neighbours 0.939 → 0.805d_b rescored +0.354 → +0.332d_a rescored -0.256 → -0.235d_a mined -0.256d_b mined 0.354min_cos_consec (site) 0.9431min_cos_anchor (site) 0.9235dataset emolialang enspeaker EN_tWXJSSWaLoytotal 17.2schain gain +3.2 dBseam step 2.0 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(formal, authoritative) Further south are the stately plantation houses owned by sugar planters, mostly standing on one of the lots in the family hacienda.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.3/10; 7.1s, EN.
EN_tWXJSSWaLoy_W000044 · in -14.7 dBFS · gain -5.3 dB · emolia-00954
(casual, formal) Inside the haciendas are chapels whose altar and icons date back to 1917.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, formal; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 2.0/10; 5.2s, EN.
EN_tWXJSSWaLoy_W000045 · in -15.8 dBFS · gain -4.2 dB · emolia-00954
(formal, authoritative) Educational visits to these places may be arranged at the BAI's City Tourism Office.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.6/10; 5.3s, EN.
EN_tWXJSSWaLoy_W000046 · in -15.3 dBFS · gain -4.7 dB · emolia-00954
Sourness ↓  /  Prideidentity −0.02 emotion 53 %   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #15

This chain comes from the two-sided rule: it only counts if both emotions move — Sourness down and Pride up — by at least 0.25 each.

The chain starts with Pride clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.30.

At the same time Sourness goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.11, then +0.19 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 39 s · no · eurospeech

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.934 before conversion and 0.914 after — it fell by 0.021. Neighbour-to-neighbour the worst pair went 0.934 → 0.914. (The earlier render, with segment 1 left raw, scores 0.807 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.298 in the original and +0.158 after conversion — 53 % of the delta retained. On the other named axis, Sourness, -0.263 became -0.021.

Quality. Mean predicted overall quality across the segments went 2.95 → 3.31 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.934 → 0.914 -0.021identity cos neighbours 0.934 → 0.914d_b rescored +0.298 → +0.158d_a rescored -0.263 → -0.021d_a mined -0.263d_b mined 0.298min_cos_consec (site) 0.9392min_cos_anchor (site) 0.9392dataset eurospeechlang nospeaker norway_9338-2total 38.8schain gain +1.2 dBseam step 0.6 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an elderly masculine voice · slightly cool, dark, slightly rough, quiet background, fairly steady, frequent disfluency, slurred, audible breath
(sourness, intoxication altered states of consciousness, concentration · measured, very low-energy, relaxed, casual) å drive en ordinær vareproduksjon i et høykostland. (surprised gasp) Det har vi sett i andre næringer. Det er enorm konkurranse om arbeidskraften, og det er kostbart å sikre inntektsutviklingen.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is slightly cool, dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, intoxication altered states of consciousness, concentration; style: casual, monologue; below-average recording, quiet background; genuineness 4.4/6; vocal-burst blend 4.2/10; 12.6s, NO.
norway_9338-2_8962496_8975056 · in -39.7 dBFS · gain +19.7 dB · eurospeech-02317
(sourness, disgust, impatience and irritability · normal-paced, normally alert, neutral tension, casual) inntektsnivået i Norge 30 pst. høyere, og nå reiser svenske ungdommer hit i hopetall. Det viser litt av utfordringen som vi har løst gjennom at vi faktisk har økt inntektene med
full caption & clip details
An elderly masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is slightly cool, dark, slightly rough, slightly thin; slurred, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, disgust, impatience and irritability; style: casual, monologue; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 5.4/10; 11.8s, NO.
norway_9338-2_8990784_9002560 · in -39.3 dBFS · gain +19.3 dB · eurospeech-02317
(pride, intoxication altered states of consciousness, thankfulness gratitude · measured, normally alert, neutral tension, casual) 120 000 kr per årsverk i de årene vi har vært i regjering, fram til vi fikk brudd nå, og det er en kvalitet ved de tiltakene og ved de oppgjørene vi har hatt tidligere. Så vi har levert – vi greide ikke det i år.
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is slightly cool, dark, slightly rough, thin; slurred, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as pride, intoxication altered states of consciousness, thankfulness gratitude; style: casual, monologue; below-average recording, quiet background; genuineness 4.2/6; vocal-burst blend 7.0/10; 14.8s, NO.
norway_9338-2_9002560_9017328 · in -39.0 dBFS · gain +19.0 dB · eurospeech-02317
Concentration ↓  /  Contemplationidentity −0.08 emotion 98 %   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #16

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Contemplation up — by at least 0.25 each.

The chain starts with Contemplation around average — 0.44, lower than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.48.

At the same time Concentration goes the other way, from 0.79 (higher than 79 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.10, then -0.01, then +0.19 — not a clean run: step 3 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 18 s · snippets

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.806 before conversion and 0.726 after — it fell by 0.080. Neighbour-to-neighbour the worst pair went 0.817 → 0.721. (The earlier render, with segment 1 left raw, scores 0.669 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.477 in the original and +0.468 after conversion — 98 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.269 became -0.379.

Quality. Mean predicted overall quality across the segments went 2.39 → 2.65 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.806 → 0.726 -0.080identity cos neighbours 0.817 → 0.721d_b rescored +0.477 → +0.468d_a rescored -0.269 → -0.379d_a mined -0.268d_b mined 0.481min_cos_consec (site) 0.8066min_cos_anchor (site) 0.7898dataset snippetslang undspeaker batch280_part1_batch280_patotal 16.9schain gain +1.9 dBseam step 2.1 dBcrossfades 100/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(some disfluency, casual, monologue) We have to look at an equation that can include the spin
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 3.4/6; vocal-burst blend 1.6/10; 3.1s.
batch280_part1_batch280_part1_chunk_993_1_770737 · in -22.9 dBFS · gain +2.9 dB · snippets-00943
(some disfluency, casual, monologue) are polymetric or alphametric are two by two matrices.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 0.9/10; 3.9s.
batch280_part1_batch280_part1_chunk_993_1_770834 · in -22.7 dBFS · gain +2.7 dB · snippets-00943
(infatuation, emotional numbness · some disfluency, casual, monologue) interaction between the electromagnetic fields and electrons.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation, emotional numbness; style: casual, monologue; good recording, no background noise; genuineness 2.7/6; vocal-burst blend 2.4/10; 3.3s.
batch280_part1_batch280_part1_chunk_993_1_770842 · in -30.0 dBFS · gain +10.0 dB · snippets-00943
(little disfluency, casual, monologue) So we have these two different ways to represent a state.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 1.3/10; 3.5s.
batch280_part1_batch280_part1_chunk_993_1_770908 · in -24.1 dBFS · gain +4.1 dB · snippets-00943
(contemplation · little disfluency, casual, monologue) of course, still without any understanding of where it comes from.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation; style: casual, monologue; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 1.2/10; 3.6s.
batch280_part1_batch280_part1_chunk_993_1_771004 · in -26.2 dBFS · gain +6.2 dB · snippets-00943
Concentration ↓  /  Thankfulness Gratitudeidentity −0.04 emotion 167 %   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #17

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Thankfulness Gratitude up — by at least 0.25 each.

The chain starts with Thankfulness Gratitude around average — 0.42, lower than 58 % of clips in this corpus — and ends with it clearly present at 0.70, higher than 70 % of clips in this corpus. That is a total rise of 0.28.

At the same time Concentration goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.55 (higher than 55 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.09, then +0.19 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 24 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.808 before conversion and 0.770 after — it fell by 0.038. Neighbour-to-neighbour the worst pair went 0.874 → 0.816. (The earlier render, with segment 1 left raw, scores 0.622 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.291 in the original and +0.486 after conversion — 167 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.308 became -0.421.

Quality. Mean predicted overall quality across the segments went 2.61 → 3.01 (+0.40) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.808 → 0.770 -0.038identity cos neighbours 0.874 → 0.816d_b rescored +0.291 → +0.486d_a rescored -0.308 → -0.421d_a mined -0.310d_b mined 0.291min_cos_consec (site) 0.8291min_cos_anchor (site) 0.8763dataset emolialang enspeaker EN_B00059_S07001total 23.5schain gain +2.1 dBseam step 0.4 dBcrossfades 100/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normally alert, slightly relaxed, fairly steady
(normal-paced, some disfluency, casual, monologue) So we had a look at the seismic that was available. You can see on this slide the light grey lines. (low mumble) A lot of seismic already collected in the north, but there were some very large and obvious gaps.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 2.6/6; vocal-burst blend 2.8/10; 11.4s, EN.
EN_B00059_S07001_W000007 · in -25.0 dBFS · gain +5.0 dB · emolia-01373
(awe, interest · normal-paced, some disfluency, monologue, whispered) This time we were looking at the layers in the basin and the seismic tells us something about the size and the thickness of those layers.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, interest; style: monologue, whispered; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 1.1/10; 7.9s, EN.
EN_B00059_S07001_W000008 · in -19.8 dBFS · gain -0.2 dB · emolia-01373
(measured, little disfluency, whispered, monologue) Similar story to the energy geochemistry. For the first time.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: whispered, monologue; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 1.9/10; 4.6s, EN.
EN_B00059_S07001_W000009 · in -19.4 dBFS · gain -0.6 dB · emolia-01373
Pride ↓  /  Doubtidentity −0.01 emotion 72 %   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #18

This chain comes from the two-sided rule: it only counts if both emotions move — Pride down and Doubt up — by at least 0.25 each.

The chain starts with Doubt around average — 0.55, higher than 55 % of clips in this corpus — and ends with it strongly present at 0.81, higher than 81 % of clips in this corpus. That is a total rise of 0.26.

At the same time Pride goes the other way, from 0.80 (higher than 80 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.04, then +0.21 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 30 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.901 before conversion and 0.892 after — it fell by 0.009. Neighbour-to-neighbour the worst pair went 0.920 → 0.861. (The earlier render, with segment 1 left raw, scores 0.851 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.256 in the original and +0.185 after conversion — 72 % of the delta retained, which is most of it. On the other named axis, Pride, -0.276 became -0.167.

Quality. Mean predicted overall quality across the segments went 3.20 → 3.31 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.901 → 0.892 -0.009identity cos neighbours 0.920 → 0.861d_b rescored +0.256 → +0.185d_a rescored -0.276 → -0.167d_a mined -0.276d_b mined 0.256min_cos_consec (site) 0.9304min_cos_anchor (site) 0.9251dataset emolialang zhspeaker ZH_B00059_S02677total 29.6schain gain +1.7 dBseam step 0.8 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, measured, normally alert
(monologue, formal) 把你的计划付诸实际行动,不要开空头支票,也不要为自己定下在一年之内收入有千倍上升这样不切实际的计划。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 1.2/10; 10.8s, ZH.
ZH_B00059_S02677_W000003 · in -18.7 dBFS · gain -1.3 dB · emolia-03869
(didactic, monologue) 你要学会循序渐进,一步一步的实现你的计划,首先控制你的债务,然后购买资产,哪怕是小小的一个。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.9/10; 11.2s, ZH.
ZH_B00059_S02677_W000004 · in -18.7 dBFS · gain -1.3 dB · emolia-03869
(whispered, ASMR) 学习如何花钱,并做出适当的改变。最后付诸行动,我们说过许多遍了。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: whispered, ASMR; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.3/10; 8.0s, ZH.
ZH_B00059_S02677_W000005 · in -19.1 dBFS · gain -0.9 dB · emolia-03869
Pride ↓  /  Emotional Numbnessidentity −0.09 emotion 62 %   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #19

This chain comes from the two-sided rule: it only counts if both emotions move — Pride down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.26.

At the same time Pride goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.09, then -0.03, then -0.03 — not a clean run: step 3 moves back the other way by 0.03 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.97 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.97), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 83 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.837 before conversion and 0.750 after — it fell by 0.087. Neighbour-to-neighbour the worst pair went 0.946 → 0.897. (The earlier render, with segment 1 left raw, scores 0.734 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.259 in the original and +0.160 after conversion — 62 % of the delta retained. On the other named axis, Pride, -0.338 became -0.307.

Quality. Mean predicted overall quality across the segments went 2.98 → 3.17 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.837 → 0.750 -0.087identity cos neighbours 0.946 → 0.897d_b rescored +0.259 → +0.160d_a rescored -0.338 → -0.307d_a mined -0.338d_b mined 0.260min_cos_consec (site) 0.9646min_cos_anchor (site) 0.9682dataset emolialang enspeaker EN_RNVju-I4_AYtotal 82.2schain gain +2.0 dBseam step 0.9 dBcrossfades 150/150/100/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, good recording, normally alert, slightly relaxed, steady, no disfluency, clear
(pride, infatuation, hope enthusiasm optimism · normal-paced, moderate pitch range, minimal breath, newsreading) It stresses the importance of science and technology for achieving sustainable and equitable socioeconomic growth and poverty eradication.Two primary policy documents operationalize the SADC Treaty of 1992, the Regional Indicative Strategic Development Plan for 2005–2020, adopted in 2003, and the Strategic Indicative Plan for the Organ
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as pride, infatuation, hope enthusiasm optimism; style: newsreading, formal; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.0/10; 25.9s, EN.
EN_RNVju-I4_AY_W000089 · in -15.3 dBFS · gain -4.7 dB · emolia-01009
(concentration, triumph · measured, moderate pitch range, no audible breath, formal) The Regional Indicative Strategic Development Plan for 2005-2020 identifies the region's 12 priority areas for both sectorial and cross-cutting intervention, mapping out goals and setting up concrete targets for each
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, no audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, triumph; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 14.9s, EN.
EN_RNVju-I4_AY_W000090 · in -15.5 dBFS · gain -4.5 dB · emolia-01009
(emotional numbness, concentration, pain · measured, fairly narrow pitch, no audible breath, formal) The four sectorial areas are, trade and economic liberalization, infrastructure, sustainable food security and human and social development. The eight cross-cutting areas are, poverty, combating the HIV, AIDS pandemic, gender equality, science and technology,
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, no audible breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, concentration, pain; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 21.9s, EN.
EN_RNVju-I4_AY_W000091 · in -16.7 dBFS · gain -3.3 dB · emolia-01009
(emotional numbness, concentration · measured, fairly narrow pitch, light breath, formal) Information and Communication Technologies' ICTs Environment and Sustainable Development Private Sector Development, and Statistics, targets include
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, concentration; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 13.1s, EN.
EN_RNVju-I4_AY_W000092 · in -16.6 dBFS · gain -3.4 dB · emolia-01009
(emotional numbness · normal-paced, moderate pitch range, light breath, formal) Ensuring that 50% of decision-making positions in the public sector are held by women by 2015
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 7.1s, EN.
EN_RNVju-I4_AY_W000093 · in -15.5 dBFS · gain -4.5 dB · emolia-01009
Pride ↓  /  Concentrationidentity −0.07 emotion 92 %   emotion_twosided__AB2__T0.25__C0.25__INTERNAL · #20

This chain comes from the two-sided rule: it only counts if both emotions move — Pride down and Concentration up — by at least 0.25 each.

The chain starts with Concentration clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.27.

At the same time Pride goes the other way, from 0.81 (higher than 81 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.17, then -0.08, then +0.18 — not a clean run: step 2 moves back the other way by 0.08 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 38 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.930 before conversion and 0.861 after — it fell by 0.069. Neighbour-to-neighbour the worst pair went 0.859 → 0.822. (The earlier render, with segment 1 left raw, scores 0.737 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.268 in the original and +0.246 after conversion — 92 % of the delta retained, which is essentially all of it. On the other named axis, Pride, -0.280 became -0.038.

Quality. Mean predicted overall quality across the segments went 2.96 → 3.07 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.930 → 0.861 -0.069identity cos neighbours 0.859 → 0.822d_b rescored +0.268 → +0.246d_a rescored -0.280 → -0.038d_a mined -0.280d_b mined 0.268min_cos_consec (site) 0.9376min_cos_anchor (site) 0.9478dataset emolialang enspeaker EN_sVncJx6pCeYtotal 36.6schain gain +0.1 dBseam step 3.2 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(steady, newsreading, formal) As a status class, the Intelligencia includes artists, teachers and academics, writers, journalists, and the literary haums de letra
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 8.9s, EN.
EN_sVncJx6pCeY_W000002 · in -16.0 dBFS · gain -4.0 dB · emolia-00663
(fear · fairly steady, formal, newsreading) The intelligentsia status class arose in the late 18th century, in Russian-controlled Poland, during the Age of Partitions 1772–95
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.6s, EN.
EN_sVncJx6pCeY_W000003 · in -14.4 dBFS · gain -5.6 dB · emolia-00663
(emotional numbness · steady, formal, authoritative) In practice, the status and social function of the intelligentsia varied by society
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.4/10; 5.7s, EN.
EN_sVncJx6pCeY_W000007 · in -15.1 dBFS · gain -4.9 dB · emolia-00663
(steady, newsreading, formal) In Eastern Europe, intellectuals were deprived of political influence and access to the effective levers of economic development.The intelligentsia were at the functional periphery of their societies
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.9s, EN.
EN_sVncJx6pCeY_W000008 · in -15.0 dBFS · gain -5.0 dB · emolia-00663