proxy_taillift__PXR__T0.25__C0.25__INTERNAL — voice-corrected

Manifest tier. proxy_taillift, rule PXR, T=0.25, step cap 0.25. Population 346,174 chains (3,818 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 278,176.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_proxy_taillift__PXR__T0.25__C0.25__INTERNAL.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
54segments re-voiced
0.748 → 0.750median worst-to-anchor identity cosine
87 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Disgust ↓  /  Disappointmentidentity −0.04 emotion 102 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #1

This chain comes from the proxy rule: the same two-sided test as above, but because Disappointment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Disappointment below average — 0.39, lower than 61 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.49.

At the same time Disgust goes the other way, from 0.71 (higher than 71 % of clips in this corpus) to 0.33 (lower than 67 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.00, then +0.49 — a slow start, with most of the change arriving in the final step.

The largest step is 0.49, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.83 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.83 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 29 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.769 before conversion and 0.732 after — it fell by 0.038. Neighbour-to-neighbour the worst pair went 0.817 → 0.714. (The earlier render, with segment 1 left raw, scores 0.687 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.494 in the original and +0.502 after conversion — 102 % of the delta retained, which is essentially all of it. On the other named axis, Disgust, -0.383 became -0.388.

Quality. Mean predicted overall quality across the segments went 2.69 → 2.95 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.769 → 0.732 -0.038identity cos neighbours 0.817 → 0.714d_b rescored +0.494 → +0.502d_a rescored -0.383 → -0.388d_a mined -0.383d_b mined 0.494min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00081_S00388total 27.5schain gain +2.3 dBseam step 0.8 dBcrossfades 150/100/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, normally alert, slightly relaxed, fairly steady
(normal-paced, some disfluency, casual, didactic) So what we did, basically in the end, was define two extension properties on the point class.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, didactic; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.7/10; 5.3s, EN.
EN_B00081_S00388_W000000 · in -18.2 dBFS · gain -1.8 dB · emolia-01802
(pride · normal-paced, frequent disfluency, monologue) So that's a (low mumble) pretty powerful concept and enables you to create (ahem) asynchronous code more easily.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride; style: monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 1.2/10; 5.8s, EN.
EN_B00081_S00388_W000001 · in -19.4 dBFS · gain -0.6 dB · emolia-01802
(measured, frequent disfluency, casual, conversational) (low mumble) Uhm, you do notice that the IDE support is not as mature as (low mumble) the support for Java yet. (low mumble) Uhm, sometimes it's a bit slow. On my machine it's a bit slower than on his machine.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 2.1/10; 11.7s, EN.
EN_B00081_S00388_W000002 · in -19.4 dBFS · gain -0.6 dB · emolia-01802
(normal-paced, frequent disfluency, conversational, casual) So it's getting better and better, but it's still not at the same level as the JavaScript.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: conversational, casual; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 1.1/10; 5.2s, EN.
EN_B00081_S00388_W000003 · in -22.4 dBFS · gain +2.4 dB · emolia-01802
Fear ↓  /  Emotional Numbnessidentity +0.02 emotion 93 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #2

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness around average — 0.57, higher than 57 % of clips in this corpus — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.27.

At the same time Fear goes the other way, from 0.78 (higher than 78 % of clips in this corpus) to 0.45 (lower than 55 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.07 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.68 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.82 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.68, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 20 s · de · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.674 before conversion and 0.689 after — it rose by 0.015. Neighbour-to-neighbour the worst pair went 0.707 → 0.758. (The earlier render, with segment 1 left raw, scores 0.677 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.272 in the original and +0.254 after conversion — 93 % of the delta retained, which is essentially all of it. On the other named axis, Fear, -0.308 became -0.101.

Quality. Mean predicted overall quality across the segments went 2.55 → 3.09 (+0.53) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.674 → 0.689 +0.015identity cos neighbours 0.707 → 0.758d_b rescored +0.272 → +0.254d_a rescored -0.308 → -0.101d_a mined -0.333d_b mined 0.269min_cos_consec (site) 0.8202min_cos_anchor (site) 0.6813dataset emolialang despeaker DE_QRUOVjKO0zEtotal 19.1schain gain +0.9 dBseam step 1.1 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, measured, slightly relaxed, steady, frequent disfluency
(normally alert, didactic, monologue) also, genau, wie das Leitgewebe-Cambium, ist auch das Cork-Cambium.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 0.0/10; 5.1s, DE.
DE_QRUOVjKO0zE_W000017 · in -15.3 dBFS · gain -4.7 dB · emolia-00193
(concentration · subdued, whispered, didactic) In beide Richtungen aktiv, es erzeugt, gegen außen Kork Fellleben und gegen innen sekundäres Rentenparenchym, das Feloderm.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: whispered, didactic; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 0.2/10; 10.9s, DE.
DE_QRUOVjKO0zE_W000018 · in -15.8 dBFS · gain -4.2 dB · emolia-00193
(normally alert, didactic, monologue) Ja, auf der Seite sehen wir diese drei
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 0.4/10; 3.4s, DE.
DE_QRUOVjKO0zE_W000019 · in -15.9 dBFS · gain -4.1 dB · emolia-00193
Pride ↓  /  Affectionidentity +0.28 emotion REVERSED   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #3

This chain comes from the proxy rule: the same two-sided test as above, but because Affection is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Affection clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.32.

At the same time Pride goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.13 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.43 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.60 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.43, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 15 s · es · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.463 before conversion and 0.742 after — it rose by 0.279. Neighbour-to-neighbour the worst pair went 0.603 → 0.819. (The earlier render, with segment 1 left raw, scores 0.521 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

The emotional move did not survive. Re-scored end to end, Affection moved +0.325 in the original and -0.200 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Pride, -0.288 became +0.102.

Quality. Mean predicted overall quality across the segments went 2.62 → 2.88 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.463 → 0.742 +0.279identity cos neighbours 0.603 → 0.819d_b rescored +0.325 → -0.200d_a rescored -0.288 → +0.102d_a mined -0.287d_b mined 0.324min_cos_consec (site) 0.6033min_cos_anchor (site) 0.4285dataset podcastlang esspeaker 162167total 14.8schain gain +2.1 dBseam step 6.5 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, fast, normally alert, some disfluency, light breath
(pride · slightly relaxed, fairly steady, average clarity, casual) todo el mundo dice que literalmente yo creo que es la mejor actuación que ha
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride; style: casual, playful; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 5.1/10; 3.4s, ES.
162167_00165456 · in -18.0 dBFS · gain -2.0 dB · podcast-04461
(slightly relaxed, fairly steady, average clarity, conversational) (surprised gasp) Él actuó super cabrón. Y supuestamente también le quitaron la como que el mérito por ser nominado a los
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational, casual; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 5.7/10; 5.8s, ES.
162167_00166784 · in -25.5 dBFS · gain +5.5 dB · podcast-04457
(affection, embarrassment · neutral tension, moderately variable, somewhat unclear, casual) Era como que, ok, estás tomando una buena decisión. Qué difícil es no tirar spoilers, ¿verdad? Estás en la tierra. Sí, por favor.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, dark, fairly smooth, thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as affection, embarrassment; style: casual, playful; below-average recording, some background noise; genuineness 6.0/6; vocal-burst blend 7.1/10; 5.9s, ES.
162167_00171752 · in -16.8 dBFS · gain -3.2 dB · podcast-04481
Fear ↓  /  Intoxication Altered States of Consciousnessidentity +0.12 emotion 46 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #4

This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Intoxication Altered States of Consciousness around average — 0.57, higher than 57 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.28.

At the same time Fear goes the other way, from 0.87 (higher than 87 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.14, then +0.14 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.73 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 14 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.744 before conversion and 0.865 after — it rose by 0.121. Neighbour-to-neighbour the worst pair went 0.707 → 0.876. (The earlier render, with segment 1 left raw, scores 0.742 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.277 in the original and +0.126 after conversion — 46 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Fear, -0.357 became +0.128.

Quality. Mean predicted overall quality across the segments went 2.67 → 2.87 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.744 → 0.865 +0.121identity cos neighbours 0.707 → 0.876d_b rescored +0.277 → +0.126d_a rescored -0.357 → +0.128d_a mined -0.357d_b mined 0.277min_cos_consec (site) 0.7282min_cos_anchor (site) 0.7856dataset emolialang zhspeaker ZH_B00016_S03252total 13.8schain gain +1.9 dBseam step 1.5 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(brisk, some disfluency, clear, authoritative) 手机上除了这个时间功能,还可以看到你在哪个APP上花的时间更多,你也可以把它分享在你的评论区里面。
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, monologue; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 2.6/10; 7.3s, ZH.
ZH_B00016_S03252_W000906 · in -21.8 dBFS · gain +1.8 dB · emolia-03434
(normal-paced, some disfluency, average clarity, formal) 因为你在哪个app上花的时间最多。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 3.4/6; vocal-burst blend 1.9/10; 3.5s, ZH.
ZH_B00016_S03252_W000907 · in -22.2 dBFS · gain +2.2 dB · emolia-03434
(measured, no disfluency, average clarity, formal) 哪个app就是创业者的元宇宙。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 2.8/6; vocal-burst blend 2.9/10; 3.4s, ZH.
ZH_B00016_S03252_W000908 · in -25.6 dBFS · gain +5.6 dB · emolia-03434
Teasing ↓  /  Astonishment Surpriseidentity +0.03 emotion 65 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #5

This chain comes from the proxy rule: the same two-sided test as above, but because Astonishment Surprise is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Astonishment Surprise clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.33.

At the same time Teasing goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.54. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.06, then +0.19, then +0.08 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 45 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.536 before conversion and 0.570 after — it rose by 0.035. Neighbour-to-neighbour the worst pair went 0.543 → 0.649. (The earlier render, with segment 1 left raw, scores 0.419 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.329 in the original and +0.213 after conversion — 65 % of the delta retained. On the other named axis, Teasing, -0.540 became -0.564.

Quality. Mean predicted overall quality across the segments went 2.65 → 2.83 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.536 → 0.570 +0.035identity cos neighbours 0.543 → 0.649d_b rescored +0.329 → +0.213d_a rescored -0.540 → -0.564d_a mined -0.540d_b mined 0.329min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00050_S07922total 43.6schain gain +2.8 dBseam step 3.3 dBcrossfades 150/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, measured, normally alert, slightly relaxed, light breath
(teasing · steady, no disfluency, clear, narration) Now and then you might see the lights of a shop or a small restaurant.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as teasing; style: narration, monologue; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 1.8/10; 3.5s, EN.
EN_B00050_S07922_W000001 · in -16.8 dBFS · gain -3.2 dB · emolia-01210
(longing, fatigue exhaustion · fairly steady, almost no disfluency, clear, narration) But most of the doors belong to business places that had been closed hours ago. Then the cops suddenly slowed his walk.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing, fatigue exhaustion; style: narration, monologue; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.0/10; 7.9s, EN.
EN_B00050_S07922_W000002 · in -17.7 dBFS · gain -2.3 dB · emolia-01210
(intoxication altered states of consciousness, relief, awe · steady, almost no disfluency, slurred, narration) Near the door of a darkened shop, a man was standing. As the cop walked toward him, the man spoke quickly. It's all right, officer, he said. I'm waiting for a friend. Twenty years ago, we agreed to meet here tonight. It sounds strange to you, doesn't it? I'll explain if you want to be sure that everything's all right. About twenty years ago, there was a restaurant here where this shop stands. Big Joe Brady's Restaurant.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, dark, slightly rough, slightly thin; slurred, almost no disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as intoxication altered states of consciousness, relief, awe; style: narration, storytelling; average recording, quiet background; genuineness 0.8/6; vocal-burst blend 1.2/10; 29.5s, EN.
EN_B00050_S07922_W000003 · in -17.2 dBFS · gain -2.8 dB · emolia-01210
(astonishment surprise, longing · fairly steady, almost no disfluency, clear, storytelling) It was here until five years ago, said the comp.
full caption & clip details
A child masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as astonishment surprise, longing; style: storytelling, narration; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.9/10; 3.2s, EN.
EN_B00050_S07922_W000004 · in -15.7 dBFS · gain -4.3 dB · emolia-01210
Interest ↓  /  Contemplationidentity +0.04 emotion 75 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #6

This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contemplation clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.26.

At the same time Interest goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.12 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 37 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.845 before conversion and 0.886 after — it rose by 0.040. Neighbour-to-neighbour the worst pair went 0.932 → 0.930. (The earlier render, with segment 1 left raw, scores 0.845 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.257 in the original and +0.192 after conversion — 75 % of the delta retained, which is most of it. On the other named axis, Interest, -0.249 became -0.308.

Quality. Mean predicted overall quality across the segments went 3.04 → 3.10 (+0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.845 → 0.886 +0.040identity cos neighbours 0.932 → 0.930d_b rescored +0.257 → +0.192d_a rescored -0.249 → -0.308d_a mined -0.250d_b mined 0.257min_cos_consec (site) 0.9397min_cos_anchor (site) 0.8882dataset podcastlang enspeaker 670896total 36.2schain gain +2.8 dBseam step 2.6 dBcrossfades 150/100 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, fairly smooth, good recording, no background noise
(interest · brisk, normally alert, slightly relaxed, casual) ancient Greek textbooks that were school textbooks. So if you were at school in the first century, uh (low mumble) you you would have been exposed to these sort of tools, and it was taught as a tool for how you how you write, you make a text come alive by making it vivid and graphic and dramatic.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, full; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as interest; style: casual, conversational; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 3.3/10; 18.0s, EN.
670896_00025528 · in -29.4 dBFS · gain +9.3 dB · podcast-01878
(emotional numbness, amusement · normal-paced, normally alert, slightly relaxed, casual) But ultimately it has a rhetorical function, which is to appeal to emotions. So that led me into a bit of a (ahem) um I guess a rabbit hole (breathy giggle)
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as emotional numbness, amusement; style: casual, playful; good recording, no background noise; genuineness 4.0/6; vocal-burst blend 0.8/10; 11.1s, EN.
670896_00027327 · in -28.4 dBFS · gain +8.4 dB · podcast-01876
(contemplation · slow, very low-energy, relaxed, casual) (ahem) in the in my PhD of thinking about how this author, John, was
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, slightly vulnerable; reads as contemplation; style: casual, conversational; good recording, no background noise; genuineness 3.0/6; vocal-burst blend 1.7/10; 7.4s, EN.
670896_00028439 · in -28.5 dBFS · gain +8.5 dB · podcast-03730
Hope Enthusiasm Optimism ↓  /  Interestidentity −0.08 emotion 122 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #7

This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Interest clearly present — 0.59, higher than 59 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.32.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.87 (higher than 87 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then -0.03, then +0.24, then -0.05 — not a clean run: step 2 moves back the other way by 0.03 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 47 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.841 before conversion and 0.764 after — it fell by 0.077. Neighbour-to-neighbour the worst pair went 0.841 → 0.793. (The earlier render, with segment 1 left raw, scores 0.693 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.322 in the original and +0.393 after conversion — 122 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Hope Enthusiasm Optimism, -0.261 became -0.084.

Quality. Mean predicted overall quality across the segments went 2.55 → 2.99 (+0.44) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.841 → 0.764 -0.077identity cos neighbours 0.841 → 0.793d_b rescored +0.322 → +0.393d_a rescored -0.261 → -0.084d_a mined -0.260d_b mined 0.321min_cos_consec (site) 0.8198min_cos_anchor (site) 0.8198dataset emolialang enspeaker EN_B00023_S01649total 46.0schain gain +4.8 dBseam step 1.6 dBcrossfades 100/150/100/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normal-paced, normally alert, slightly relaxed
(frequent disfluency, casual, playful) So we hit that and it's gonna take us to alibaba.com. It's gonna hit the search. (low mumble)
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, playful; good recording, quiet background; genuineness 4.6/6; vocal-burst blend 2.1/10; 5.1s, EN.
EN_B00023_S01649_W000031 · in -16.1 dBFS · gain -3.9 dB · emolia-00704
(hope enthusiasm optimism, elation · some disfluency, casual, monologue) This brings up 975 suppliers, far too many to navigate through. So we need to filter that right down in order to find the absolute best suppliers for this.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism, elation; style: casual, monologue; good recording, no background noise; genuineness 2.7/6; vocal-burst blend 3.7/10; 8.2s, EN.
EN_B00023_S01649_W000032 · in -16.8 dBFS · gain -3.2 dB · emolia-00704
(some disfluency, casual, monologue) So, uh, (ahem) in a way to bring this down, we want to select, (ahem) uh, trade assurance, as we mentioned before, so our payment is protected.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.8/6; vocal-burst blend 1.7/10; 6.1s, EN.
EN_B00023_S01649_W000033 · in -15.9 dBFS · gain -4.1 dB · emolia-00704
(triumph, interest, concentration · some disfluency, casual, monologue) And we want to select verified to make sure that everything they say is true.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, interest, concentration; style: casual, monologue; good recording, quiet background; genuineness 4.0/6; vocal-burst blend 3.6/10; 13.9s, EN.
EN_B00023_S01649_W000034 · in -16.5 dBFS · gain -3.5 dB · emolia-00704
(interest, infatuation · some disfluency, casual, monologue) (ahem) Uh, for the blocking on the lens for the US market, they already know that because they supply this market and they'll already have that testing carried out if they're a top supplier. So as you select North America and Western Europe, now we get it down to 235.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, infatuation; style: casual, monologue; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 7.2/10; 13.3s, EN.
EN_B00023_S01649_W000035 · in -16.7 dBFS · gain -3.3 dB · emolia-00704
Intoxication Altered States of Consciousness ↓  /  Infatuationidentity +0.66 emotion REVERSED   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #8

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation around average — 0.51, right about the corpus median — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.34.

At the same time Intoxication Altered States of Consciousness goes the other way, from 0.83 (higher than 83 % of clips in this corpus) to 0.49 (right about the corpus median), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.09, then +0.24 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.07 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.12 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.07, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 15 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.024 before conversion and 0.689 after — it rose by 0.665. Neighbour-to-neighbour the worst pair went 0.100 → 0.677. (The earlier render, with segment 1 left raw, scores 0.591 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

The emotional move did not survive. Re-scored end to end, Infatuation moved +0.336 in the original and -0.608 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Intoxication Altered States of Consciousness, -0.343 became -0.558.

Quality. Mean predicted overall quality across the segments went 2.75 → 2.89 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.024 → 0.689 +0.665identity cos neighbours 0.100 → 0.677d_b rescored +0.336 → -0.608d_a rescored -0.343 → -0.558d_a mined -0.343d_b mined 0.336min_cos_consec (site) 0.1231min_cos_anchor (site) 0.0703dataset podcastlang enspeaker 887330total 14.0schain gain +2.5 dBseam step 2.1 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert, slightly relaxed
(frequent disfluency, average clarity, wide pitch range, casual) can last for generations. That's I mean, that's a powerful concept.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 3.0/6; vocal-burst blend 2.0/10; 4.3s, EN.
887330_00099400 · in -24.8 dBFS · gain +4.8 dB · podcast-03792
(interest, contemplation · little disfluency, clear, moderate pitch range, casual) What what are some of the key principles that Nash highlights when it comes to building that kind of legacy?
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as interest, contemplation; style: casual, monologue; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.6/10; 5.5s, EN.
887330_00099936 · in -26.8 dBFS · gain +6.8 dB · podcast-03802
(almost no disfluency, average clarity, moderate pitch range, casual) It all boils down to two main things. Really starting early and being consistent.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 2.1/10; 4.5s, EN.
887330_00100536 · in -24.7 dBFS · gain +4.7 dB · podcast-03781
Shame ↓  /  Impatience and Irritabilityidentity −0.05 emotion 95 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #9

This chain comes from the proxy rule: the same two-sided test as above, but because Impatience and Irritability is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Impatience and Irritability clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.25.

At the same time Shame goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are -0.02, then +0.05, then +0.23 — not a clean run: step 1 moves back the other way by 0.02 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 93 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.855 before conversion and 0.801 after — it fell by 0.054. Neighbour-to-neighbour the worst pair went 0.878 → 0.881. (The earlier render, with segment 1 left raw, scores 0.691 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.252 in the original and +0.239 after conversion — 95 % of the delta retained, which is essentially all of it. On the other named axis, Shame, -0.294 became -0.331.

Quality. Mean predicted overall quality across the segments went 2.75 → 3.30 (+0.55) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.855 → 0.801 -0.054identity cos neighbours 0.878 → 0.881d_b rescored +0.252 → +0.239d_a rescored -0.294 → -0.331d_a mined -0.294d_b mined 0.251min_cos_consec (site) 0.8941min_cos_anchor (site) 0.8601dataset podcastlang enspeaker 416437total 91.7schain gain +6.2 dBseam step 1.4 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · average recording, quiet background, neutral tension, moderately variable, frequent disfluency
(shame, pain, helplessness · normal-paced, normally alert, average clarity, casual) And secondly, it was just that my mind kept going back to the person I'm with now. My mind just kept going back to wow, this is the man I want to be with for the rest of my life. (low mumble) Um
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, neutral openness; reads as shame, pain, helplessness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.4/6; vocal-burst blend 3.6/10; 15.2s, EN.
416437_00075252 · in -29.1 dBFS · gain +9.1 dB · podcast-02754
(infatuation, shame, affection · measured, very low-energy, somewhat unclear, casual) he asked me last night, and I'm not gonna tell y'all, but he said, Are you gonna marry me? I said, Yes, I am 100%. I am. I am 10 toes down for this man. Y'all don't understand how much I love this man, like it is unbelievable to my like my soul just is like, Wow, okay, like sometimes I'd be like, Okay,
full caption & clip details
A child somewhat masculine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as infatuation, shame, affection; style: casual, monologue; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 9.1/10; 29.9s, EN.
416437_00076772 · in -29.3 dBFS · gain +9.3 dB · podcast-02793
(doubt, contemplation, triumph · measured, normally alert, average clarity, casual) are we doing good? Are we fighting? Are we doing this? Are we doing that? Are we in between? What are we doing? What are we doing? And sometimes it's good to step back a little bit, collectively, get yourself together, and then come back on that topic, or come back onto whatever you and your partner need to talk about later.
full caption & clip details
A child masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as doubt, contemplation, triumph; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 8.7/10; 23.0s, EN.
416437_00079760 · in -25.5 dBFS · gain +5.5 dB · podcast-06249
(impatience and irritability, sourness, contempt · normal-paced, energised, somewhat unclear, casual) Because in the moment that person's probably pissed off, and you're making it you're just throwing it into the flames, like you're just you're making the fire worse, like it's just getting worse. You can't have two hot heads just trying to get their point across because one's trying to be funny and one's other just trying to be serious, and then somebody's gonna have a mental breakdown.
full caption & clip details
A young adult somewhat masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as impatience and irritability, sourness, contempt; style: casual, monologue; average recording, quiet background; mildly explicit content; genuineness 4.3/6; vocal-burst blend 9.5/10; 24.0s, EN.
416437_00082064 · in -27.2 dBFS · gain +7.2 dB · podcast-02753
Interest ↓  /  Embarrassmentidentity +0.19 emotion 51 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #10

This chain comes from the proxy rule: the same two-sided test as above, but because Embarrassment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Embarrassment clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.26.

At the same time Interest goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.46. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.03, then +0.22, then -0.09, then +0.09 — not a clean run: step 3 moves back the other way by 0.09 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.58 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.62 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.58, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 58 s · en · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.559 before conversion and 0.753 after — it rose by 0.195. Neighbour-to-neighbour the worst pair went 0.597 → 0.753. (The earlier render, with segment 1 left raw, scores 0.574 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.258 in the original and +0.133 after conversion — 51 % of the delta retained. On the other named axis, Interest, -0.461 became -0.445.

Quality. Mean predicted overall quality across the segments went 2.58 → 3.10 (+0.52) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.559 → 0.753 +0.195identity cos neighbours 0.597 → 0.753d_b rescored +0.258 → +0.133d_a rescored -0.461 → -0.445d_a mined -0.456d_b mined 0.256min_cos_consec (site) 0.6154min_cos_anchor (site) 0.5802dataset podcastlang enspeaker 585546total 56.7schain gain +2.8 dBseam step 2.4 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, normal-paced, normally alert
(interest, astonishment surprise, amusement · neutral tension, moderately variable, frequent disfluency, casual) one. Anyway, he did an episode with Jay Z, which was fuck which was lit, but in it, he went to Rick Rubin's (low mumble) uh recording studio. And hung out with this girl, I think her name's Madison Ward, but she sings the song Mirrors. Oh
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, astonishment surprise, amusement; style: casual, conversational; average recording, some background noise; genuineness 5.8/6; vocal-burst blend 9.0/10; 14.7s, EN.
585546_00094136 · in -30.9 dBFS · gain +10.9 dB · podcast-04043
(awe, pleasure ecstasy, infatuation · fully relaxed, fairly steady, frequent disfluency, casual) Yeah. It was all it's really beautiful. She it's like mixed, it's like Nora Jones mixed with like Alicia Keys. It's
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as awe, pleasure ecstasy, infatuation; style: casual, monologue; average recording, quiet background; mildly explicit content; genuineness 3.5/6; vocal-burst blend 4.1/10; 6.9s, EN.
585546_00095648 · in -30.2 dBFS · gain +10.2 dB · podcast-04141
(amusement, pleasure ecstasy, interest · neutral tension, moderately variable, some disfluency, casual) Yeah, my bad bomb, coconut. (exhausted groan) Um just listen to it. She's in a she I looked up, looked her up. She's like a volleyball player. Uh she just graduated college, she's like 6'4, and she would sing at parties, and people were like, dude, you should try to sing it. So she's like, Alright, so she moved to Cali. So check her out. Our
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as amusement, pleasure ecstasy, interest; style: casual, conversational; good recording, quiet background; genuineness 4.7/6; vocal-burst blend 7.4/10; 15.6s, EN.
585546_00096759 · in -30.6 dBFS · gain +10.6 dB · podcast-04025
(elation, sexual lust, infatuation · neutral tension, fairly steady, some disfluency, casual) Alright. (ahem) Uh my last song then, since he's doing something to do too. I don't even need to explain this. It's just probably one of like honestly, probably one of the greatest songs that have ever been written, and it's just smash mouth. It's all star.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, slightly thin; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as elation, sexual lust, infatuation; style: casual, conversational; below-average recording, quiet background; genuineness 5.3/6; vocal-burst blend 10.0/10; 11.3s, EN.
585546_00098744 · in -29.8 dBFS · gain +9.8 dB · podcast-04012
(embarrassment · neutral tension, fairly steady, some disfluency, casual) No, wait, wait, well, let's say Dalton, explain what the quote of the week is. Oh, okay. So the (low mumble) quote of the week is uh what we're gonna do is we're gonna tell you guys a quote that we like and then just not explain it.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as embarrassment; style: casual, conversational; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 5.1/10; 8.8s, EN.
585546_00100783 · in -26.1 dBFS · gain +6.1 dB · podcast-04130
Pride ↓  /  Reliefidentity +0.13 emotion 29 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #11

This chain comes from the proxy rule: the same two-sided test as above, but because Relief is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Relief around average — 0.45, lower than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.53.

At the same time Pride goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.45. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.21, then +0.20, then +0.12 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.78 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.75 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.78, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 76 s · lv · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.752 before conversion and 0.878 after — it rose by 0.126. Neighbour-to-neighbour the worst pair went 0.724 → 0.869. (The earlier render, with segment 1 left raw, scores 0.606 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.534 in the original and +0.155 after conversion — 29 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Pride, -0.451 became -0.314.

Quality. Mean predicted overall quality across the segments went 2.85 → 3.52 (+0.67) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.752 → 0.878 +0.126identity cos neighbours 0.724 → 0.869d_b rescored +0.534 → +0.155d_a rescored -0.451 → -0.314d_a mined -0.451d_b mined 0.534min_cos_consec (site) 0.7495min_cos_anchor (site) 0.7824dataset eurospeechlang lvspeaker latvia_20140122115419total 75.2schain gain +2.6 dBseam step 0.8 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · slightly rough, balanced body, average recording, quiet background, normally alert, normal breath
(pride, anger, contempt · normal-paced, slightly relaxed, fairly steady, authoritative) mums ir spēcīga tieslietu sistēma, valstī nav korupcijas, korupcijas apkarotāji cīnās viens ar otru, un valstī vairs arī nav ko dalīt.” Nepārprotami, šī valdība būs nākamās turpinājums, un ir skaidrs, ka par iedzīvotājiem atkal netiks domāts.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, almost no disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as pride, anger, contempt; style: authoritative, cartoonish; average recording, quiet background; genuineness 0.5/6; vocal-burst blend 1.0/10; 18.7s, LV.
latvia_20140122115419_1227008_1245712 · in -13.8 dBFS · gain -6.2 dB · eurospeech-01998
(contempt, bitterness, malevolence malice · measured, slightly relaxed, fairly steady, cartoonish) Šī gada budžets arī netiks grozīts, bet, plānojot nākamā gada budžetu, kā pirmo prioritāti atkal ieraudzīsim tankus. Tanki vispār ir zīmīgs simbols šai koalīcijai. Arī izdaudzināto ekonomikas stabilizāciju izjutīs vienPrudentiaīpašnieki, nevis ierindas iedzīvotāji,
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as contempt, bitterness, malevolence malice; style: cartoonish, authoritative; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 2.1/10; 19.8s, LV.
latvia_20140122115419_1245712_1265536 · in -14.8 dBFS · gain -5.2 dB · eurospeech-01998
(bitterness, contempt, triumph · normal-paced, neutral tension, moderately variable, cartoonish) kuriem arī šogad par apkuri būs jāmaksā lieli rēķini, jo koalīcija negrib, piemēram, samazināt pievienotās vērtības nodokli apkurei. Cilvēki atkal nevarēs saņemt pienācīgu medicīnisko aprūpi, bet veselības ministre televīzijā vien atzīs: „Jā, tas tā ir, un tur neko vairs nevar darīt.”
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as bitterness, contempt, triumph; style: cartoonish, monologue; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 2.3/10; 19.5s, LV.
latvia_20140122115419_1265536_1285072 · in -15.9 dBFS · gain -4.1 dB · eurospeech-01998
(relief · measured, slightly relaxed, fairly steady, cartoonish) Bet ir arī pozitīvā nots - šīs koalīcijas pastāvēšanas ilgums. Tautai tā būs jāpiedzīvo vien astoņus mēnešus, pēc kuriem situāciju mēs visi varēsim arī labot. Paldies.(SC frakcijas deputātu aplausi.)
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as relief; style: cartoonish, authoritative; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.3/10; 17.8s, LV.
latvia_20140122115419_1285072_1302855 · in -16.2 dBFS · gain -3.8 dB · eurospeech-01998
Infatuation ↓  /  Confusionidentity +0.02 emotion 52 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #12

This chain comes from the proxy rule: the same two-sided test as above, but because Confusion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Confusion clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.29.

At the same time Infatuation goes the other way, from 0.83 (higher than 83 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.15, then +0.16, then -0.02 — not a clean run: step 3 moves back the other way by 0.02 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.78 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.78 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.78, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 28 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.662 before conversion and 0.686 after — it rose by 0.024. Neighbour-to-neighbour the worst pair went 0.662 → 0.719. (The earlier render, with segment 1 left raw, scores 0.701 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.289 in the original and +0.150 after conversion — 52 % of the delta retained. On the other named axis, Infatuation, -0.297 became +0.004.

Quality. Mean predicted overall quality across the segments went 2.94 → 3.16 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.662 → 0.686 +0.024identity cos neighbours 0.662 → 0.719d_b rescored +0.289 → +0.150d_a rescored -0.297 → +0.004d_a mined -0.297d_b mined 0.292min_cos_consec (site) 0.7837min_cos_anchor (site) 0.7837dataset emolialang zhspeaker ZH_B00041_S02467total 27.2schain gain +1.7 dBseam step 0.5 dBcrossfades 150/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · fairly smooth, balanced body, no background noise, measured, steady, clear, fairly narrow pitch, light breath
(very low-energy, slightly relaxed, no disfluency, whispered) 实际上,此时你关注的是疾病。
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, slightly relaxed, steady; timbre is neutral-toned, dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: whispered, formal; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 2.8/10; 3.1s, ZH.
ZH_B00041_S02467_W000002 · in -17.2 dBFS · gain -2.8 dB · emolia-03689
(emotional numbness · subdued, slightly relaxed, no disfluency, whispered) 你的注意力的所在将会吸引他的本质。当你说我想要钱,但我终究得不到。
full caption & clip details
A young adult feminine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is slightly warm, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; reads as emotional numbness; style: whispered, ASMR; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.3/10; 9.1s, ZH.
ZH_B00041_S02467_W000003 · in -17.4 dBFS · gain -2.6 dB · emolia-03689
(emotional numbness, confusion, thankfulness gratitude · very low-energy, relaxed, no disfluency, whispered) 此时,你就是在对匮乏的关注,你等于在说来吧,没钱的生活。
full caption & clip details
An elderly feminine voice; delivery is very low-energy, measured, relaxed, steady; timbre is slightly warm, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, submissive, neutral openness; reads as emotional numbness, confusion, thankfulness gratitude; style: whispered, ASMR; average recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.6/10; 8.7s, ZH.
ZH_B00041_S02467_W000004 · in -18.4 dBFS · gain -1.6 dB · emolia-03689
(confusion, fear · very low-energy, relaxed, frequent disfluency, whispered) 当你以吸引其接近的方式来思考金钱时,你总会心情愉悦。
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, relaxed, steady; timbre is slightly warm, slightly dark, fairly smooth, balanced body; clear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, submissive, neutral openness; reads as confusion, fear; style: whispered, ASMR; average recording, no background noise; genuineness 1.7/6; vocal-burst blend 1.3/10; 6.8s, ZH.
ZH_B00041_S02467_W000005 · in -16.4 dBFS · gain -3.6 dB · emolia-03689
Hope Enthusiasm Optimism ↓  /  Contentmentidentity +0.51 emotion 127 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #13

This chain comes from the proxy rule: the same two-sided test as above, but because Contentment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contentment around average — 0.57, higher than 57 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.34.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.19 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.22 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.15 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.22, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 28 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.234 before conversion and 0.745 after — it rose by 0.511. Neighbour-to-neighbour the worst pair went 0.198 → 0.730. (The earlier render, with segment 1 left raw, scores 0.749 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.335 in the original and +0.425 after conversion — 127 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Hope Enthusiasm Optimism, -0.301 became -0.316.

Quality. Mean predicted overall quality across the segments went 2.77 → 3.00 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.234 → 0.745 +0.511identity cos neighbours 0.198 → 0.730d_b rescored +0.335 → +0.425d_a rescored -0.301 → -0.316d_a mined -0.300d_b mined 0.342min_cos_consec (site) 0.1499min_cos_anchor (site) 0.2193dataset podcastlang enspeaker 937386total 27.5schain gain -0.0 dBseam step 1.1 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · fairly smooth, good recording, no background noise, moderate pitch range
(hope enthusiasm optimism · measured, very low-energy, slightly relaxed, whispered) if you're applying to a big international companies, they would ask. Yeah.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, slightly relaxed, steady; timbre is slightly warm, neutral-bright, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, audible breath; affect is mildly negative, slightly submissive, neutral openness; reads as hope enthusiasm optimism; style: whispered, storytelling; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 0.7/10; 5.8s, EN.
937386_00115448 · in -29.1 dBFS · gain +9.1 dB · podcast-04348
(pride · normal-paced, normally alert, slightly relaxed, whispered) No, I wouldn't say start to. For example, my company is international. They did ask me for the degree which I provided.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as pride; style: whispered, narration; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 0.3/10; 6.3s, EN.
937386_00116448 · in -29.5 dBFS · gain +9.5 dB · podcast-04368
(contentment, relief · normal-paced, normally alert, neutral tension, casual) So no, basically I have a senior role and I don't have a degree. And but it also is I was flexible with my visa, right? So it's (low mumble) uh it's also not on the visa that I have a senior role. (low mumble) Um so it also depends on what you want to get out of (low mumble) uh your your agreement with
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contentment, relief; style: casual, conversational; good recording, no background noise; genuineness 4.3/6; vocal-burst blend 7.4/10; 15.6s, EN.
937386_00117119 · in -25.4 dBFS · gain +5.4 dB · podcast-04337
Fear ↓  /  Interestidentity +0.01 emotion 87 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #14

This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Interest clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.28.

At the same time Fear goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.58 (higher than 58 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.12, then +0.16 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 26 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.844 before conversion and 0.855 after — it rose by 0.011. Neighbour-to-neighbour the worst pair went 0.844 → 0.847. (The earlier render, with segment 1 left raw, scores 0.789 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.276 in the original and +0.241 after conversion — 87 % of the delta retained, which is most of it. On the other named axis, Fear, -0.356 became +0.028.

Quality. Mean predicted overall quality across the segments went 3.02 → 3.08 (+0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.844 → 0.855 +0.011identity cos neighbours 0.844 → 0.847d_b rescored +0.276 → +0.241d_a rescored -0.356 → +0.028d_a mined -0.356d_b mined 0.276min_cos_consec (site) 0.8393min_cos_anchor (site) 0.8393dataset emolialang enspeaker EN_oyIVSZ_nSEgtotal 25.8schain gain +2.5 dBseam step 2.1 dBcrossfades 100/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady, some disfluency
(fear · normal breath, casual, monologue) If you're in group A, you will also be coming for week two. That's when you're gonna need to have your personal protective equipment. So, that's going to include
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as fear; style: casual, monologue; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 0.0/10; 10.3s, EN.
EN_oyIVSZ_nSEg_W000034 · in -18.8 dBFS · gain -1.2 dB · emolia-01803
(pain · light breath, casual, monologue) Goggles that go all the way around your face, not ones that look like sunglasses. They have to, they have to actually touch your face right here. Okay? And,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as pain; style: casual, monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.6/10; 9.0s, EN.
EN_oyIVSZ_nSEg_W000035 · in -18.6 dBFS · gain -1.4 dB · emolia-01803
(interest · light breath, casual, monologue) You're going to want to look for some with vents because we will also have masks on. Masking is mandatory all over campus. (low mumble)
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as interest; style: casual, monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 1.3/10; 6.9s, EN.
EN_oyIVSZ_nSEg_W000036 · in -17.9 dBFS · gain -2.1 dB · emolia-01803
Concentration ↓  /  Infatuationidentity −0.01 emotion 87 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #15

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation around average — 0.54, higher than 54 % of clips in this corpus — and ends with it strongly present at 0.81, higher than 81 % of clips in this corpus. That is a total rise of 0.28.

At the same time Concentration goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.43 (lower than 57 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.14, then -0.02, then +0.16 — not a clean run: step 2 moves back the other way by 0.02 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 19 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.881 before conversion and 0.869 after — it fell by 0.012. Neighbour-to-neighbour the worst pair went 0.876 → 0.869. (The earlier render, with segment 1 left raw, scores 0.797 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.305 in the original and +0.265 after conversion — 87 % of the delta retained, which is most of it. On the other named axis, Concentration, -0.286 became -0.218.

Quality. Mean predicted overall quality across the segments went 2.98 → 3.01 (+0.03) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.881 → 0.869 -0.012identity cos neighbours 0.876 → 0.869d_b rescored +0.305 → +0.265d_a rescored -0.286 → -0.218d_a mined -0.286d_b mined 0.276min_cos_consec (site) 0.8773min_cos_anchor (site) 0.9136dataset emolialang zhspeaker ZH_B00033_S08202total 17.8schain gain -1.1 dBseam step 0.4 dBcrossfades 100/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, measured, normally alert
(steady, formal, monologue) 由此会产生一种破坏欲,如果有机会,哪怕自己过不好,也不能让对方得意。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.8/10; 6.3s, ZH.
ZH_B00033_S08202_W000071 · in -16.8 dBFS · gain -3.2 dB · emolia-03604
(fairly steady, narration, monologue) 若是赵耳有这样的想法,那么他就会进侯府为妾。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 3.0/10; 4.0s, ZH.
ZH_B00033_S08202_W000072 · in -16.7 dBFS · gain -3.3 dB · emolia-03604
(steady, formal, monologue) 至少他没有想要鱼死网破。至少他心中有光。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 3.0/10; 3.9s, ZH.
ZH_B00033_S08202_W000073 · in -16.1 dBFS · gain -3.9 dB · emolia-03604
(steady, formal, monologue) 才能选择成全对方的明媚,成全自己的下半生。
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 1.4/10; 4.1s, ZH.
ZH_B00033_S08202_W000074 · in -15.9 dBFS · gain -4.1 dB · emolia-03604
Concentration ↓  /  Longingidentity −0.08 emotion 27 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #16

This chain comes from the proxy rule: the same two-sided test as above, but because Longing is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Longing around average — 0.56, higher than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.40.

At the same time Concentration goes the other way, from 0.92 (higher than 92 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.19, then -0.01, then +0.22 — not a clean run: step 2 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 33 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.823 before conversion and 0.747 after — it fell by 0.076. Neighbour-to-neighbour the worst pair went 0.793 → 0.759. (The earlier render, with segment 1 left raw, scores 0.691 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.406 in the original and +0.109 after conversion — 27 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Concentration, -0.253 became -0.160.

Quality. Mean predicted overall quality across the segments went 2.85 → 3.02 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.823 → 0.747 -0.076identity cos neighbours 0.793 → 0.759d_b rescored +0.406 → +0.109d_a rescored -0.253 → -0.160d_a mined -0.252d_b mined 0.400min_cos_consec (site) 0.8862min_cos_anchor (site) 0.9126dataset emolialang zhspeaker ZH_B00007_S00700total 32.3schain gain +2.3 dBseam step 2.1 dBcrossfades 150/100/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, average recording, no background noise, normally alert, slightly relaxed, fairly steady, no disfluency
(concentration · measured, moderate pitch range, monologue, whispered) 二是农村二三产业发展缓慢,农村非农业产业就业岗位增长受限。三是农村劳动力素质整体偏低。必要的职业技能培训还未完全跟上。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, whispered; average recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.8/10; 12.5s, ZH.
ZH_B00007_S00700_W000047 · in -22.3 dBFS · gain +2.3 dB · emolia-03342
(pain · measured, moderate pitch range, monologue, whispered) 四是管理体制工作机制和工作手段未完全建立。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as pain; style: monologue, whispered; average recording, no background noise; genuineness 1.5/6; vocal-burst blend 2.2/10; 4.3s, ZH.
ZH_B00007_S00700_W000048 · in -23.3 dBFS · gain +3.3 dB · emolia-03342
(fast, moderate pitch range, monologue, formal) 和谐社会的关键是让老百姓自己感觉到社会的和谐。
full caption & clip details
A young adult feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 1.5/6; vocal-burst blend 4.1/10; 4.1s, ZH.
ZH_B00007_S00700_W000049 · in -21.3 dBFS · gain +1.3 dB · emolia-03342
(longing, thankfulness gratitude, pain · measured, fairly narrow pitch, monologue, whispered) 构建社会主义和谐社会,必须解决好买房贵、上学、贵、看病贵、就业难等新的民生问题。从古到今,老有所养病,有所依。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing, thankfulness gratitude, pain; style: monologue, whispered; average recording, no background noise; genuineness 1.4/6; vocal-burst blend 1.0/10; 11.9s, ZH.
ZH_B00007_S00700_W000050 · in -21.8 dBFS · gain +1.8 dB · emolia-03342
Emotional Numbness ↓  /  Impatience and Irritabilityidentity +0.04 emotion 95 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #17

This chain comes from the proxy rule: the same two-sided test as above, but because Impatience and Irritability is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Impatience and Irritability clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.33.

At the same time Emotional Numbness goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.51 (higher than 51 % of clips in this corpus), a change of -0.44. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.20, then -0.06, then +0.19 — not a clean run: step 2 moves back the other way by 0.06 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 57 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.639 before conversion and 0.676 after — it rose by 0.037. Neighbour-to-neighbour the worst pair went 0.842 → 0.821. (The earlier render, with segment 1 left raw, scores 0.565 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.326 in the original and +0.311 after conversion — 95 % of the delta retained, which is essentially all of it. On the other named axis, Emotional Numbness, -0.437 became -0.190.

Quality. Mean predicted overall quality across the segments went 2.92 → 3.19 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.639 → 0.676 +0.037identity cos neighbours 0.842 → 0.821d_b rescored +0.326 → +0.311d_a rescored -0.437 → -0.190d_a mined -0.437d_b mined 0.326min_cos_consec (site) 0.8779min_cos_anchor (site) 0.8489dataset emolialang enspeaker EN_YXMLGjnzoYytotal 55.9schain gain +3.5 dBseam step 0.6 dBcrossfades 150/100/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, average clarity, moderate pitch range
(emotional numbness · slightly relaxed, fairly steady, some disfluency, formal) Some would suggest moving to the Hare Clark model, and I no doubt expect to get some questions on this (low mumble) afterwards.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, authoritative; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 1.0/10; 4.9s, EN.
EN_YXMLGjnzoYy_W000170 · in -21.8 dBFS · gain +1.8 dB · emolia-00459
(contempt, doubt, intoxication altered states of consciousness · slightly relaxed, fairly steady, little disfluency, conversational) I don't agree with this proposal as there is an issue of scale with Hare Clark. Hare Clark may work well for electing the lower house of parliament in Tasmania in the ACT.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contempt, doubt, intoxication altered states of consciousness; style: conversational, authoritative; good recording, quiet background; genuineness 1.9/6; vocal-burst blend 2.7/10; 12.5s, EN.
EN_YXMLGjnzoYy_W000171 · in -22.2 dBFS · gain +2.2 dB · emolia-00459
(concentration, triumph · slightly relaxed, fairly steady, some disfluency, monologue) when the quota is around 10,000 voters. But there is a vast difference of scale in electing the New South Wales Senate with a quota of 600,000 when that election is being conducted at the same time as the lower house of parliament. For all, if the only reason for introducing Robson rotation is to break the control of parties to control the order of election of candidates, then I think it's an artifice which isn't, isn't required. There are other ways to do it. And anyway,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, triumph; style: monologue, formal; good recording, quiet background; genuineness 1.1/6; vocal-burst blend 1.0/10; 23.7s, EN.
EN_YXMLGjnzoYy_W000172 · in -23.3 dBFS · gain +3.3 dB · emolia-00459
(impatience and irritability, anger, bitterness · neutral tension, moderately variable, some disfluency, casual) It's pretty clear and has for a long time that people are voting on party in the Senate. And just because people are voting for parties and you think they should be voting for candidates, I don't see is that a good reason to stop people from voting for parties in that way. I believe the Senate electoral system has been operating
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as impatience and irritability, anger, bitterness; style: casual, monologue; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 6.3/10; 15.3s, EN.
EN_YXMLGjnzoYy_W000173 · in -24.6 dBFS · gain +4.5 dB · emolia-00459
Impatience and Irritability ↓  /  Emotional Numbnessidentity −0.13 emotion 102 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #18

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.25.

At the same time Impatience and Irritability goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.09, then -0.04, then +0.20 — not a clean run: step 2 moves back the other way by 0.04 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.75 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 70 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.585 before conversion and 0.458 after — it fell by 0.127. Neighbour-to-neighbour the worst pair went 0.524 → 0.415. (The earlier render, with segment 1 left raw, scores 0.412 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.253 in the original and +0.259 after conversion — 102 % of the delta retained, which is essentially all of it. On the other named axis, Impatience and Irritability, -0.344 became -0.399.

Quality. Mean predicted overall quality across the segments went 3.09 → 3.19 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.585 → 0.458 -0.127identity cos neighbours 0.524 → 0.415d_b rescored +0.253 → +0.259d_a rescored -0.344 → -0.399d_a mined -0.336d_b mined 0.253min_cos_consec (site) 0.7523min_cos_anchor (site) 0.8113dataset podcastlang enspeaker 440220total 69.0schain gain +4.0 dBseam step 2.6 dBcrossfades 150/100/150 ms
Script — 4 chunks, 4 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, fairly steady, moderate pitch range, light breath
(impatience and irritability, astonishment surprise, helplessness · measured, subdued, slightly relaxed, casual) Like they even had to sign. Like there's proof that they signed this paper before they did the shows that they weren't gonna do this, that, and the other. (low mumble) Um I'm looking through the article right now. I was trying to see if if (low mumble) uh like they weren't allowed to remove clothing, they weren't allowed to use provocative speech, they weren't allowed to, you know, different things like that. And (low mumble) uh
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as impatience and irritability, astonishment surprise, helplessness; style: casual, monologue; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 7.1/10; 23.0s, EN.
440220_00625152 · in -29.9 dBFS · gain +9.9 dB · podcast-06131
(intoxication altered states of consciousness, confusion · normal-paced, normally alert, neutral tension, casual) (ahem) no use of alcohol or smoking, like all the like real traditional, you know, things that they just were like and they agreed to it, and then they get up on stage and I guess they played a completely different set list than the one
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as intoxication altered states of consciousness, confusion; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.5/6; vocal-burst blend 8.5/10; 13.8s, EN.
440220_00627452 · in -29.4 dBFS · gain +9.3 dB · podcast-06142
(infatuation, disgust, intoxication altered states of consciousness · measured, subdued, relaxed, casual) yeah, real weird. (low mumble) Um don't like that. But (ahem) uh that was apparently his protest to their anti LGBTQ laws and he basically said I wish I could find the quote, but he was like (low mumble) uh something about and they were real ticked because they said he was swearing, he was drinking, he was smoking, doing all the things they told him not to on stage. But anyway, yeah, that was that was a little while ago. It wasn't like the other day, but
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as infatuation, disgust, intoxication altered states of consciousness; style: casual, monologue; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 7.2/10; 29.4s, EN.
440220_00629912 · in -28.8 dBFS · gain +8.8 dB · podcast-06114
(normal-paced, normally alert, slightly relaxed, casual) (low mumble) They uh they were being sued before two point four million
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, no background noise; genuineness 4.0/6; vocal-burst blend 2.7/10; 3.3s, EN.
440220_00633104 · in -27.9 dBFS · gain +7.9 dB · podcast-02343
Pride ↓  /  Painidentity +0.05 emotion 142 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #19

This chain comes from the proxy rule: the same two-sided test as above, but because Pain is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Pain below average — 0.35, lower than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.56.

At the same time Pride goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.17, then +0.19 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.84 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 43 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.779 before conversion and 0.832 after — it rose by 0.053. Neighbour-to-neighbour the worst pair went 0.823 → 0.836. (The earlier render, with segment 1 left raw, scores 0.602 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.562 in the original and +0.796 after conversion — 142 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Pride, -0.688 became -0.699.

Quality. Mean predicted overall quality across the segments went 3.09 → 3.18 (+0.09) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.779 → 0.832 +0.053identity cos neighbours 0.823 → 0.836d_b rescored +0.562 → +0.796d_a rescored -0.688 → -0.699d_a mined -0.402d_b mined 0.562min_cos_consec (site) 0.8360min_cos_anchor (site) 0.7862dataset emolialang enspeaker EN_m6ev4IZOaaItotal 42.2schain gain +2.7 dBseam step 1.0 dBcrossfades 150/100/150 ms
Script — 4 chunks, 4 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, balanced body, quiet background, fairly steady
(pride · normal-paced, normally alert, slightly relaxed, casual) Yeah, as we reported, they're on track to deliver 500,000 EVs this year, which is a significant amount. That's way ahead of (low mumble) everybody else except for Tesla. Uh, (low mumble) Herbert Deese was their CEO that put all of this in motion.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as pride; style: casual, conversational; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.4/10; 13.9s, EN.
EN_m6ev4IZOaaI_W000255 · in -22.8 dBFS · gain +2.8 dB · emolia-01586
(infatuation, awe, hope enthusiasm optimism · normal-paced, normally alert, slightly relaxed, casual) And (low mumble) uhm, you know, he really had a radical vision for VW and really felt like it had to be a radical remaking of the company.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as infatuation, awe, hope enthusiasm optimism; style: casual, conversational; good recording, quiet background; genuineness 4.0/6; vocal-burst blend 8.1/10; 7.8s, EN.
EN_m6ev4IZOaaI_W000256 · in -24.4 dBFS · gain +4.5 dB · emolia-01586
(normal-paced, normally alert, slightly relaxed, casual) Or, you know, they were gonna run into problems and, (low mumble) uh, so yeaah, so he started a lot of ambitious programs that have gotten them to 500,000 EVs a year, which is significant.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; good recording, quiet background; genuineness 3.4/6; vocal-burst blend 4.3/10; 9.9s, EN.
EN_m6ev4IZOaaI_W000257 · in -22.6 dBFS · gain +2.6 dB · emolia-01586
(pain · slow, very low-energy, relaxed, conversational) But (low mumble) uhm, he was sort of moved out recently as CEO and the new CEO is definitely scaling back these plans, (low mumble) uhm, to be much much less ambitious.
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as pain; style: conversational, casual; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 2.7/10; 11.2s, EN.
EN_m6ev4IZOaaI_W000258 · in -21.6 dBFS · gain +1.6 dB · emolia-01586
Emotional Numbness ↓  /  Contemptidentity −0.02 emotion 104 %   proxy_taillift__PXR__T0.25__C0.25__INTERNAL · #20

This chain comes from the proxy rule: the same two-sided test as above, but because Contempt is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contempt clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.26.

At the same time Emotional Numbness goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.67 (higher than 67 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.05, then +0.22 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.97 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 50 s · da · eurospeech

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.930 before conversion and 0.915 after — it fell by 0.015. Neighbour-to-neighbour the worst pair went 0.930 → 0.915. (The earlier render, with segment 1 left raw, scores 0.803 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contempt moved +0.262 in the original and +0.271 after conversion — 104 % of the delta retained, which is essentially all of it. On the other named axis, Emotional Numbness, -0.311 became -0.240.

Quality. Mean predicted overall quality across the segments went 3.12 → 3.49 (+0.37) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.930 → 0.915 -0.015identity cos neighbours 0.930 → 0.915d_b rescored +0.262 → +0.271d_a rescored -0.311 → -0.240d_a mined -0.310d_b mined 0.263min_cos_consec (site) 0.9666min_cos_anchor (site) 0.9600dataset eurospeechlang daspeaker denmark_20161M049_2017-01-total 49.4schain gain +2.8 dBseam step 0.2 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-bright, slightly rough, balanced body, average recording, quiet background, measured, normally alert, wide pitch range
(emotional numbness, disgust, concentration · slightly relaxed, fairly steady, frequent disfluency, didactic) mig. DSB er jo en selvstændig virksomhed. DSB styrer sig selv – under tilsyn af transportministeren, ja, men DSB styrer sig selv.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, balanced body; very clear, frequent disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as emotional numbness, disgust, concentration; style: didactic, authoritative; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 0.1/10; 11.4s, DA.
denmark_20161M049_2017-01-25_1300_12385504_12396912 · in -28.0 dBFS · gain +8.0 dB · eurospeech-00284
(pride, bitterness, anger · slightly relaxed, moderately variable, almost no disfluency, authoritative) DSB, at der er problemer her, afsat 100 mio. kr. til ekstra vedligehold af tog. Det har ingen betydning for DSB's beslutning om det, at man har besluttet at tage 300 mio. kr. ud af DSB. For DSB har et overskud hvert år, der er større end de 300 mio. kr.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, almost no disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as pride, bitterness, anger; style: authoritative, monologue; average recording, quiet background; genuineness 0.8/6; vocal-burst blend 1.6/10; 19.8s, DA.
denmark_20161M049_2017-01-25_1300_12396912_12416726 · in -25.1 dBFS · gain +5.1 dB · eurospeech-00284
(contempt, disgust, concentration · neutral tension, moderately variable, frequent disfluency, authoritative) Så der er stadig overskud i DSB, selv om regeringen har taget 300 mio. kr. derfra. Det kan altså ikke begrunde, at man ikke har hensat de penge, der skal til til vedligehold af tog. Det er en fejl i DSB. Det har DSB beklaget, og det er jeg glad for at DSB har beklaget.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as contempt, disgust, concentration; style: authoritative, cartoonish; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 2.9/10; 18.6s, DA.
denmark_20161M049_2017-01-25_1300_12416726_12435296 · in -25.8 dBFS · gain +5.8 dB · eurospeech-00284