proxy_taillift__PXR__T0.40__C0.25__INTERNAL — voice-corrected

Manifest tier. proxy_taillift, rule PXR, T=0.4, step cap 0.25. Population 12,397 chains (139 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 10,431.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_proxy_taillift__PXR__T0.40__C0.25__INTERNAL.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
61segments re-voiced
0.767 → 0.757median worst-to-anchor identity cosine
81 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Contentment ↓  /  Emotional Numbnessidentity +0.01 emotion 76 %   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #1

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness around average — 0.49, right about the corpus median — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.43.

At the same time Contentment goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.50 (right about the corpus median), a change of -0.45. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are -0.05, then +0.17, then +0.14, then +0.16 — not a clean run: step 1 moves back the other way by 0.05 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.76 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.76 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.76, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 44 s · snippets

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.689 before conversion and 0.701 after — it rose by 0.012. Neighbour-to-neighbour the worst pair went 0.633 → 0.586. (The earlier render, with segment 1 left raw, scores 0.474 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.409 in the original and +0.312 after conversion — 76 % of the delta retained, which is most of it. On the other named axis, Contentment, -0.438 became -0.722.

Quality. Mean predicted overall quality across the segments went 2.85 → 3.06 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.689 → 0.701 +0.012identity cos neighbours 0.633 → 0.586d_b rescored +0.409 → +0.312d_a rescored -0.438 → -0.722d_a mined -0.445d_b mined 0.432min_cos_consec (site) 0.7596min_cos_anchor (site) 0.7596dataset snippetslang undspeaker batch30_part1_batch30_parttotal 42.6schain gain +4.2 dBseam step 1.3 dBcrossfades 100/100/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-bright, balanced body, some disfluency, average clarity, light breath
(contentment, jealousy and envy · brisk, energised, neutral tension, casual) So while still working at Louis Vuitton, he quietly banded together a group of forward-thinking creatives including his brother Garam and a handful of other designers he'd met over the years.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contentment, jealousy and envy; style: casual, storytelling; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.6/10; 11.2s.
batch30_part1_batch30_part1_chunk_1268_1_1353713 · in -35.0 dBFS · gain +14.9 dB · snippets-01048
(teasing, hope enthusiasm optimism, infatuation · brisk, energised, slightly relaxed, casual) Hey, fine, if you guys like French fashion so much, I'll just make a French brand.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as teasing, hope enthusiasm optimism, infatuation; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 3.5/10; 4.3s.
batch30_part1_batch30_part1_chunk_1268_1_1353752 · in -33.7 dBFS · gain +13.7 dB · snippets-01048
(longing, interest, affection · normal-paced, normally alert, slightly relaxed, casual) Back then, Demna and Kanye couldn't have possibly imagined the ways in which their paths would someday cross. And don't worry, we will be talking about that later. But in the meantime, Kanye's mind was simply set on discovering the next big thing in fashion.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as longing, interest, affection; style: casual, playful; good recording, quiet background; genuineness 3.0/6; vocal-burst blend 3.4/10; 13.8s.
batch30_part1_batch30_part1_chunk_1268_1_1353811 · in -34.1 dBFS · gain +14.1 dB · snippets-01048
(brisk, energised, neutral tension, casual) Once again, this collection relied on many of the principal design elements established by the first collection, but this time around he pushed them to the absolute limit.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 5.1/10; 8.5s.
batch30_part1_batch30_part1_chunk_1268_1_1353909 · in -35.3 dBFS · gain +15.3 dB · snippets-01048
(emotional numbness · brisk, normally alert, slightly relaxed, casual) So even though he didn't win, being named one of the finalists provided him with a ton of exposure.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness; style: casual, conversational; good recording, no background noise; genuineness 3.5/6; vocal-burst blend 3.2/10; 5.4s.
batch30_part1_batch30_part1_chunk_1268_1_1353993 · in -34.5 dBFS · gain +14.5 dB · snippets-01048
Contempt ↓  /  Infatuationidentity −0.04 emotion 75 %   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #2

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation below average — 0.39, lower than 61 % of clips in this corpus — and ends with it strongly present at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.44.

At the same time Contempt goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.44. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.20 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 26 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.887 before conversion and 0.844 after — it fell by 0.043. Neighbour-to-neighbour the worst pair went 0.841 → 0.822. (The earlier render, with segment 1 left raw, scores 0.767 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.436 in the original and +0.328 after conversion — 75 % of the delta retained, which is most of it. On the other named axis, Contempt, -0.440 became -0.003.

Quality. Mean predicted overall quality across the segments went 2.94 → 3.20 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.887 → 0.844 -0.043identity cos neighbours 0.841 → 0.822d_b rescored +0.436 → +0.328d_a rescored -0.440 → -0.003d_a mined -0.440d_b mined 0.436min_cos_consec (site) 0.8496min_cos_anchor (site) 0.9011dataset emolialang enspeaker EN_qc9GUj0t8O8total 25.8schain gain +1.9 dBseam step 1.0 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, good recording, brisk, energised, neutral tension, moderately variable
(contempt, sourness, disgust · some disfluency, casual, dramatic) Out of an electrified broth of not living nutrients. The Miller-Urey experiment supported the idea that all life on Earth arose in a primordial soup of basic nutrients.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contempt, sourness, disgust; style: casual, dramatic; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 1.5/10; 10.3s, EN.
EN_qc9GUj0t8O8_W000037 · in -19.0 dBFS · gain -1.0 dB · emolia-02525
(awe, astonishment surprise · some disfluency, casual, storytelling) Some scientists, though, including Crick, found this unlikely and thought life on Earth probably came from outer space.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as awe, astonishment surprise; style: casual, storytelling; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 4.1/10; 7.3s, EN.
EN_qc9GUj0t8O8_W000038 · in -19.4 dBFS · gain -0.6 dB · emolia-02525
(little disfluency, conversational, dramatic) An idea called panspermia. The discoveries of 1953 marked a new era in biology. Evolution now had a molecular basis.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; no dominant emotion; style: conversational, dramatic; good recording, quiet background; genuineness 0.8/6; vocal-burst blend 1.7/10; 8.5s, EN.
EN_qc9GUj0t8O8_W000039 · in -20.8 dBFS · gain +0.8 dB · emolia-02525
Pain ↓  /  Fatigue Exhaustionidentity −0.11 emotion 46 %   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #3

This chain comes from the proxy rule: the same two-sided test as above, but because Fatigue Exhaustion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fatigue Exhaustion below average — 0.31, lower than 69 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.52.

At the same time Pain goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.20 (lower than 80 % of clips in this corpus), a change of -0.52. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.18, then +0.23, then +0.11 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 20 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.846 before conversion and 0.735 after — it fell by 0.111. Neighbour-to-neighbour the worst pair went 0.891 → 0.814. (The earlier render, with segment 1 left raw, scores 0.502 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.524 in the original and +0.242 after conversion — 46 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Pain, -0.520 became +0.196.

Quality. Mean predicted overall quality across the segments went 2.74 → 2.81 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.846 → 0.735 -0.111identity cos neighbours 0.891 → 0.814d_b rescored +0.524 → +0.242d_a rescored -0.520 → +0.196d_a mined -0.520d_b mined 0.524min_cos_consec (site) 0.9433min_cos_anchor (site) 0.9333dataset emolialang enspeaker EN_QEoAxVsfo64total 19.3schain gain +0.8 dBseam step 0.6 dBcrossfades 100/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(formal, authoritative) Joshi continued his training with Sawai Gonda RVA
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 4.2/10; 3.0s, EN.
EN_QEoAxVsfo64_W000030 · in -13.6 dBFS · gain -6.4 dB · emolia-01760
(formal, casual) Joshi first performed live in 1941 at the age of 19
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, casual; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 3.1/10; 4.3s, EN.
EN_QEoAxVsfo64_W000032 · in -13.9 dBFS · gain -6.1 dB · emolia-01760
(formal, narration) His debut album, containing a few devotional songs in Marathi and Hindi, was released by HMV the next year in 1942
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.8s, EN.
EN_QEoAxVsfo64_W000033 · in -14.3 dBFS · gain -5.7 dB · emolia-01760
(formal, narration) Later Joshi moved to Mumbai in 1943 and worked as a radio artist
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.7/10; 4.7s, EN.
EN_QEoAxVsfo64_W000034 · in -14.8 dBFS · gain -5.2 dB · emolia-01760
Concentration ↓  /  Intoxication Altered States of Consciousnessidentity −0.07 emotion 59 %   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #4

This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Intoxication Altered States of Consciousness below average — 0.36, lower than 64 % of clips in this corpus — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.47.

At the same time Concentration goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.56 (higher than 56 % of clips in this corpus), a change of -0.43. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 26 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.799 before conversion and 0.733 after — it fell by 0.065. Neighbour-to-neighbour the worst pair went 0.854 → 0.759. (The earlier render, with segment 1 left raw, scores 0.635 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.470 in the original and +0.277 after conversion — 59 % of the delta retained. On the other named axis, Concentration, -0.430 became -0.448.

Quality. Mean predicted overall quality across the segments went 2.79 → 3.01 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.799 → 0.733 -0.065identity cos neighbours 0.854 → 0.759d_b rescored +0.470 → +0.277d_a rescored -0.430 → -0.448d_a mined -0.430d_b mined 0.467min_cos_consec (site) 0.9030min_cos_anchor (site) 0.9266dataset emolialang enspeaker EN_B00044_S01748total 25.0schain gain +1.8 dBseam step 1.0 dBcrossfades 100/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, fairly steady, clear
(concentration, interest · normal-paced, some disfluency, didactic, casual) Now images are read only, which means that once you've created an image, it cannot be changed. If you need to change something about the image, then you need to instead create a brand new image to incorporate that change.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, full; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration, interest; style: didactic, casual; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.9/10; 13.7s, EN.
EN_B00044_S01748_W000004 · in -19.2 dBFS · gain -0.8 dB · emolia-01104
(normal-paced, little disfluency, monologue, dramatic) So again, images are like blueprints for containers and they include every single thing that our application might need to run.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: monologue, dramatic; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 2.2/10; 7.4s, EN.
EN_B00044_S01748_W000005 · in -18.8 dBFS · gain -1.2 dB · emolia-01104
(measured, almost no disfluency, formal, monologue) Now containers are runnable instances of those images.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 1.4/10; 4.3s, EN.
EN_B00044_S01748_W000006 · in -18.8 dBFS · gain -1.2 dB · emolia-01104
Thankfulness Gratitude ↓  /  Concentrationidentity +0.02 emotion 124 %   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #5

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration around average — 0.49, right about the corpus median — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.48.

At the same time Thankfulness Gratitude goes the other way, from 0.88 (higher than 88 % of clips in this corpus) to 0.42 (lower than 58 % of clips in this corpus), a change of -0.46. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.18, then +0.07 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 52 s · en · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.828 before conversion and 0.850 after — it rose by 0.022. Neighbour-to-neighbour the worst pair went 0.853 → 0.879. (The earlier render, with segment 1 left raw, scores 0.753 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.482 in the original and +0.599 after conversion — 124 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Thankfulness Gratitude, -0.461 became -0.218.

Quality. Mean predicted overall quality across the segments went 2.88 → 3.24 (+0.36) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.828 → 0.850 +0.022identity cos neighbours 0.853 → 0.879d_b rescored +0.482 → +0.599d_a rescored -0.461 → -0.218d_a mined -0.461d_b mined 0.482min_cos_consec (site) 0.9330min_cos_anchor (site) 0.9146dataset eurospeechlang enspeaker uk_uk_6_25112020total 51.0schain gain +1.9 dBseam step 0.7 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-toned, slightly rough, balanced body, quiet background, measured, normally alert, slightly relaxed, fairly steady
(little disfluency, fairly narrow pitch, formal, narration) Looking beyond the ICGS, a new Member services team has also been established to provide human resources support for MPs and their staff. I should add that more than 4,000 people in Parliament
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; clear, little disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 0.1/10; 11.4s, EN.
uk_uk_6_25112020_14669168_14680576 · in -25.4 dBFS · gain +5.4 dB · eurospeech-01060
(little disfluency, moderate pitch range, narration, formal) have now taken the Valuing Everyone training, which aims to demonstrate how to recognise and understand what harassment and sexual harassment mean in the workplace and how to tackle them.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, quiet background; genuineness 1.3/6; vocal-burst blend 0.0/10; 10.1s, EN.
uk_uk_6_25112020_14680576_14690649 · in -26.7 dBFS · gain +6.7 dB · eurospeech-01060
(concentration · almost no disfluency, fairly narrow pitch, formal, newsreading) Turning to the independent expert panel, it is important to note that the appointments that we are discussing today form part of our fulfilment of the key recommendations made by Dame Laura Cox in her 2018 report.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: formal, newsreading; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 0.1/10; 11.8s, EN.
uk_uk_6_25112020_14690649_14702400 · in -26.3 dBFS · gain +6.3 dB · eurospeech-01060
(concentration · almost no disfluency, moderate pitch range, newsreading, formal) Members will remember that Dame Laura made three fundamental recommendations: the first was that Parliament’s existing policies relating to bullying, harassment or sexual harassment should be abandoned; the second was that the ICGS should be accessible to those with complaints involvinghistoricalallegations.Bothof thoserecommendations
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: newsreading, formal; average recording, quiet background; genuineness 0.2/6; vocal-burst blend 0.0/10; 18.4s, EN.
uk_uk_6_25112020_14702400_14720768 · in -25.5 dBFS · gain +5.5 dB · eurospeech-01060
Emotional Numbness ↓  /  Painidentity +0.08 emotion REVERSED   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #6

This chain comes from the proxy rule: the same two-sided test as above, but because Pain is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Pain barely there — 0.11, lower than 89 % of clips in this corpus — and ends with it clearly present at 0.72, higher than 72 % of clips in this corpus. That is a total rise of 0.61.

At the same time Emotional Numbness goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.39 (lower than 61 % of clips in this corpus), a change of -0.52. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.20, then +0.17, then +0.00 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 38 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.794 before conversion and 0.876 after — it rose by 0.082. Neighbour-to-neighbour the worst pair went 0.795 → 0.867. (The earlier render, with segment 1 left raw, scores 0.728 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Pain moved +0.609 in the original and -0.502 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Emotional Numbness, -0.517 became -0.352.

Quality. Mean predicted overall quality across the segments went 3.11 → 3.14 (+0.03) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.794 → 0.876 +0.082identity cos neighbours 0.795 → 0.867d_b rescored +0.609 → -0.502d_a rescored -0.517 → -0.352d_a mined -0.517d_b mined 0.609min_cos_consec (site) 0.8771min_cos_anchor (site) 0.8374dataset emolialang zhspeaker ZH_B00000_S09712total 36.8schain gain +1.0 dBseam step 2.4 dBcrossfades 150/150/100/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady, no disfluency
(emotional numbness · measured, narration, monologue) 如果你免疫力很差,那你只能依赖医院,只能寄希望于更高的医疗水平。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: narration, monologue; average recording, no background noise; genuineness 1.4/6; vocal-burst blend 3.9/10; 6.4s, ZH.
ZH_B00000_S09712_W000026 · in -19.6 dBFS · gain -0.5 dB · emolia-03267
(measured, narration, monologue) 我们的身体有着强大的自我修复能力,只有在平时善待他,关键时刻,他才会跟我们同舟共济。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 2.4/10; 8.8s, ZH.
ZH_B00000_S09712_W000027 · in -19.5 dBFS · gain -0.5 dB · emolia-03267
(measured, narration, monologue) 你看,人们有个头疼脑热的,自然想到上医院跑药店。可事实上,人体其实具有你想象不到的强大的自愈力。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; average recording, no background noise; genuineness 0.6/6; vocal-burst blend 3.9/10; 9.9s, ZH.
ZH_B00000_S09712_W000028 · in -20.1 dBFS · gain +0.1 dB · emolia-03267
(measured, narration, monologue) 在没有外力帮助的情况下,也能让很多疾病低下头来,这种自愈力是一种生命的本能。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 4.5/10; 8.3s, ZH.
ZH_B00000_S09712_W000029 · in -19.7 dBFS · gain -0.3 dB · emolia-03267
(normal-paced, formal, narration) 卫生部健康教育的首席专家洪昭光教授说啊。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.3/10; 4.1s, ZH.
ZH_B00000_S09712_W000030 · in -19.5 dBFS · gain -0.5 dB · emolia-03267
Interest ↓  /  Concentrationidentity −0.08 emotion 97 %   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #7

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration at the very top of the corpus — 0.95, higher than 95 % of clips in this corpus — and works its way down to around average at 0.54, higher than 54 % of clips in this corpus. That is a total fall of 0.42.

At the same time Interest goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.48 (lower than 52 % of clips in this corpus), a change of -0.51. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are -0.02, then -0.15, then -0.24 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 33 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.801 before conversion and 0.717 after — it fell by 0.084. Neighbour-to-neighbour the worst pair went 0.871 → 0.790. (The earlier render, with segment 1 left raw, scores 0.448 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved -0.415 in the original and -0.403 after conversion — 97 % of the delta retained, which is essentially all of it. On the other named axis, Interest, -0.506 became -0.440.

Quality. Mean predicted overall quality across the segments went 2.94 → 3.08 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.801 → 0.717 -0.084identity cos neighbours 0.871 → 0.790d_b rescored -0.415 → -0.403d_a rescored -0.506 → -0.440d_a mined -0.506d_b mined -0.415min_cos_consec (site) 0.8800min_cos_anchor (site) 0.9382dataset emolialang enspeaker EN_B00056_S04738total 32.2schain gain +1.3 dBseam step 1.0 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normal-paced, normally alert, slightly relaxed
(interest, sexual lust, concentration · clear, monologue, dramatic) We are driven, meaning we have mechanisms in our brain that make us motivated to pursue more of what brings both a taste of sweetness, but also that brings actual changes in blood glucose levels.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as interest, sexual lust, concentration; style: monologue, dramatic; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.5/10; 12.6s, EN.
EN_B00056_S04738_W000377 · in -15.1 dBFS · gain -4.9 dB · emolia-01330
(concentration, hope enthusiasm optimism · clear, casual, monologue) Okay, so we are motivated to eat sweet things, not just because they taste good, but because they change our blood sugar level. They increase our blood sugar level.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration, hope enthusiasm optimism; style: casual, monologue; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 3.1/10; 7.9s, EN.
EN_B00056_S04738_W000378 · in -16.7 dBFS · gain -3.3 dB · emolia-01330
(clear, monologue, didactic) This is important because it needn't be the case. It could have been that we were just wired to pursue things that taste good.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: monologue, didactic; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.3/10; 7.0s, EN.
EN_B00056_S04738_W000379 · in -15.1 dBFS · gain -4.9 dB · emolia-01330
(average clarity, casual, monologue) When the small lab, it's not a small lab, it's actually a big lab, but when Dana Small's lab,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, neutral openness; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 3.4/10; 5.3s, EN.
EN_B00056_S04738_W000380 · in -16.8 dBFS · gain -3.2 dB · emolia-01330
Malevolence Malice ↓  /  Triumphidentity −0.02 emotion 125 %   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #8

This chain comes from the proxy rule: the same two-sided test as above, but because Triumph is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Triumph below average — 0.38, lower than 62 % of clips in this corpus — and ends with it strongly present at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.42.

At the same time Malevolence Malice goes the other way, from 0.88 (higher than 88 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.51. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.22, then +0.20 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 33 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.877 before conversion and 0.861 after — it fell by 0.016. Neighbour-to-neighbour the worst pair went 0.780 → 0.789. (The earlier render, with segment 1 left raw, scores 0.788 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Triumph moved +0.418 in the original and +0.522 after conversion — 125 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Malevolence Malice, -0.513 became -0.370.

Quality. Mean predicted overall quality across the segments went 3.09 → 3.13 (+0.04) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.877 → 0.861 -0.016identity cos neighbours 0.780 → 0.789d_b rescored +0.418 → +0.522d_a rescored -0.513 → -0.370d_a mined -0.514d_b mined 0.418min_cos_consec (site) 0.8915min_cos_anchor (site) 0.9027dataset emolialang zhspeaker ZH_B00020_S00879total 31.6schain gain +1.4 dBseam step 2.0 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, measured, normally alert, slightly relaxed, no disfluency
(steady, moderate pitch range, formal, narration) 凡是能贮存的食品,他们都存在仓库里等待买主找上门来,以好价钱卖掉。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 3.0/10; 6.6s, ZH.
ZH_B00020_S00879_W000039 · in -22.6 dBFS · gain +2.6 dB · emolia-03481
(steady, moderate pitch range, formal, monologue) 因此不久就出现了一种新职业。囤积居奇者。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 4.7/10; 4.6s, ZH.
ZH_B00020_S00879_W000040 · in -22.3 dBFS · gain +2.3 dB · emolia-03481
(pride, concentration · steady, fairly narrow pitch, monologue, narration) 有些无职业的男人,带着一两个背包到农民那里挨家挨户收购食品,甚至乘火车到那些特别有利可图的地方非法收购,然后拿到城里以四五倍的价格出售。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, concentration; style: monologue, narration; average recording, no background noise; genuineness 0.6/6; vocal-burst blend 3.8/10; 13.1s, ZH.
ZH_B00020_S00879_W000041 · in -23.2 dBFS · gain +3.2 dB · emolia-03481
(fairly steady, moderate pitch range, monologue, formal) 开始农民很高兴,他们用鸡蛋和黄油换来了那么多的钞票,像流水般淌到自己的家门。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 3.4/10; 8.0s, ZH.
ZH_B00020_S00879_W000042 · in -22.4 dBFS · gain +2.4 dB · emolia-03481
Confusion ↓  /  Malevolence Maliceidentity +0.03 emotion 141 %   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #9

This chain comes from the proxy rule: the same two-sided test as above, but because Malevolence Malice is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Malevolence Malice around average — 0.49, lower than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.50.

At the same time Confusion goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.49 (lower than 51 % of clips in this corpus), a change of -0.47. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.10, then +0.24, then +0.17 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.80 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 34 s · de · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.820 before conversion and 0.846 after — it rose by 0.026. Neighbour-to-neighbour the worst pair went 0.745 → 0.796. (The earlier render, with segment 1 left raw, scores 0.815 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Malevolence Malice moved +0.504 in the original and +0.712 after conversion — 141 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Confusion, -0.471 became -0.080.

Quality. Mean predicted overall quality across the segments went 3.02 → 3.23 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.820 → 0.846 +0.026identity cos neighbours 0.745 → 0.796d_b rescored +0.504 → +0.712d_a rescored -0.471 → -0.080d_a mined -0.471d_b mined 0.504min_cos_consec (site) 0.7977min_cos_anchor (site) 0.8598dataset emolialang despeaker DE_p7ETU5g7ujutotal 33.1schain gain +3.2 dBseam step 0.8 dBcrossfades 100/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, some disfluency
(confusion · normal-paced, normally alert, slightly relaxed, conversational) Ich schätze mal, das war dann die Passvorteingabe oder so. Gut, das war jetzt mein Fehler. So, kommen wir hier zurück?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as confusion; style: conversational, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 0.0/10; 6.8s, DE.
DE_p7ETU5g7uju_W000041 · in -19.6 dBFS · gain -0.4 dB · emolia-00206
(normal-paced, normally alert, slightly relaxed, playful) Also im Prinzip simuliert es einen Spielecomputer. Ah, da komm ich zurück, okay. Was haben wir hier? Groovy Sachen.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: playful, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 0.0/10; 7.9s, DE.
DE_p7ETU5g7uju_W000042 · in -21.1 dBFS · gain +1.1 dB · emolia-00206
(disgust, teasing, sourness · normal-paced, normally alert, slightly relaxed, conversational) Ist gleich viel entspannter, oder? Also ich hab ja bei euch die Musik etwas runtergedreht.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as disgust, teasing, sourness; style: conversational, casual; average recording, quiet background; mildly explicit content; genuineness 4.5/6; vocal-burst blend 0.0/10; 4.5s, DE.
DE_p7ETU5g7uju_W000043 · in -22.0 dBFS · gain +2.0 dB · emolia-00206
(malevolence malice, anger, amusement · brisk, energised, neutral tension, cartoonish) Das war's. Wir sind dabei, um Dr. Evil zu besiegen, indem wir auch vor einer Windows-Plagiarz sitzen und irgendwelche Minispiele machen. Rockstar. Ich hatte Vertrauen in euch.
full caption & clip details
A child masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very clear, some disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as malevolence malice, anger, amusement; style: cartoonish, storytelling; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.0/10; 14.3s, DE.
DE_p7ETU5g7uju_W000044 · in -14.7 dBFS · gain -5.3 dB · emolia-00206
Interest ↓  /  Emotional Numbnessidentity +0.64 emotion 93 %   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #10

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness around average — 0.51, higher than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.43.

At the same time Interest goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.59 (higher than 59 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.10 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.16 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.10, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 52 s · de · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.131 before conversion and 0.771 after — it rose by 0.640. Neighbour-to-neighbour the worst pair went 0.160 → 0.823. (The earlier render, with segment 1 left raw, scores 0.587 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.435 in the original and +0.403 after conversion — 93 % of the delta retained, which is essentially all of it. On the other named axis, Interest, -0.411 became -0.430.

Quality. Mean predicted overall quality across the segments went 2.92 → 3.12 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.131 → 0.771 +0.640identity cos neighbours 0.160 → 0.823d_b rescored +0.435 → +0.403d_a rescored -0.411 → -0.430d_a mined -0.407d_b mined 0.434min_cos_consec (site) 0.1563min_cos_anchor (site) 0.0964dataset podcastlang despeaker 833968total 51.3schain gain +3.7 dBseam step 0.1 dBcrossfades 100/100 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert, slightly relaxed, fairly steady
(interest, contemplation, concentration · normal-paced, some disfluency, monologue, conversational) Und das ist natürlich auch was, was alle interessiert. Wie schaffe ich das? Auf der einen Seite (ahem) ganz viele Daten zu bekommen, um das, was Sie eben beschrieben haben, zu realisieren, auf der anderen Seite aber auch (low mumble) mit den Daten, den Persönlichkeitsrechten und den Daten der Bürgerinnen so umzugehen und (ahem) da nicht in Datenschutz einzugreifen. (low mumble)
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as interest, contemplation, concentration; style: monologue, conversational; good recording, quiet background; genuineness 2.6/6; vocal-burst blend 1.3/10; 29.6s, DE.
833968_00155728 · in -18.2 dBFS · gain -1.8 dB · podcast-03013
(measured, frequent disfluency, monologue, didactic) (ahem) Seine (low mumble) Daten im Netz zu hinterlegen.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.0/10; 12.1s, DE.
833968_00168441 · in -22.1 dBFS · gain +2.1 dB · podcast-02979
(emotional numbness, concentration · normal-paced, some disfluency, monologue, formal) Aber (low mumble) wie gesagt, es ist die Verantwortung des Einzelnen, die man ihm dann auch nicht abnehmen kann. Ein für den Bedarf, wie gesagt, sind (ahem) anonyme Daten völlig ausreichend.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, concentration; style: monologue, formal; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.0/10; 9.9s, DE.
833968_00169728 · in -21.0 dBFS · gain +1.0 dB · podcast-01356
Impatience and Irritability ↓  /  Elationidentity +0.50 emotion 62 %   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #11

This chain comes from the proxy rule: the same two-sided test as above, but because Elation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Elation around average — 0.46, lower than 54 % of clips in this corpus — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.42.

At the same time Impatience and Irritability goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.46. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.00, then +0.23, then +0.01, then +0.19 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.16 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.15 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.16, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 35 s · en · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.178 before conversion and 0.673 after — it rose by 0.495. Neighbour-to-neighbour the worst pair went 0.144 → 0.686. (The earlier render, with segment 1 left raw, scores 0.622 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Elation moved +0.754 in the original and +0.469 after conversion — 62 % of the delta retained. On the other named axis, Impatience and Irritability, -0.468 became -0.371.

Quality. Mean predicted overall quality across the segments went 2.66 → 2.94 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.178 → 0.673 +0.495identity cos neighbours 0.144 → 0.686d_b rescored +0.754 → +0.469d_a rescored -0.468 → -0.371d_a mined -0.460d_b mined 0.422min_cos_consec (site) 0.1491min_cos_anchor (site) 0.1619dataset podcastlang enspeaker 158097total 33.9schain gain +2.9 dBseam step 0.7 dBcrossfades 100/100/100/100 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, some disfluency, average clarity, light breath
(impatience and irritability, embarrassment, shame · brisk, energised, neutral tension, casual) what's actually going on. Sometimes they like throw it on one person if they you know, like if this player didn't do this or this one should have done that, or
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as impatience and irritability, embarrassment, shame; style: casual, playful; average recording, quiet background; mildly explicit content; genuineness 5.1/6; vocal-burst blend 8.5/10; 5.7s, EN.
158097_00254143 · in -17.3 dBFS · gain -2.7 dB · podcast-00426
(embarrassment, teasing, confusion · brisk, normally alert, neutral tension, casual) I'm I'm not even where you're at. And I just as a as a fan, I just see that and go,
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as embarrassment, teasing, confusion; style: casual, conversational; good recording, no background noise; genuineness 3.3/6; vocal-burst blend 1.7/10; 3.9s, EN.
158097_00254887 · in -22.1 dBFS · gain +2.1 dB · podcast-05319
(amusement, fatigue exhaustion, embarrassment · normal-paced, normally alert, relaxed, conversational) Oh, that that must be harsh. That must be hard to deal with. There is a lot, and and you've got to turn it off. (low mumble) Um I liken it to when you go camping and there's a mozzie in the tent. (ahem) That's just a bit annoying. Yeah. I love
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as amusement, fatigue exhaustion, embarrassment; style: conversational, casual; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 3.2/10; 10.4s, EN.
158097_00255287 · in -21.7 dBFS · gain +1.7 dB · podcast-05259
(amusement, teasing, infatuation · normal-paced, normally alert, slightly relaxed, casual) (chuckle) your analogies. Yeah, I just think that's what it is. It's just a bit annoying. And sometimes it becomes a mozzie bite, and do you scratch it or not? Oh, I'll just gonna let that one go.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly submissive, neutral openness; reads as amusement, teasing, infatuation; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.8/6; vocal-burst blend 3.0/10; 6.6s, EN.
158097_00256343 · in -18.9 dBFS · gain -1.1 dB · podcast-05271
(normal-paced, normally alert, slightly relaxed, casual) But then that yeah, that's the reality of what we're in. Now, we're all talking love, connection, unconditional, 44 sons, all that emotion and and
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 2.3/10; 7.7s, EN.
158097_00258887 · in -22.0 dBFS · gain +2.0 dB · podcast-05265
Infatuation ↓  /  Concentrationidentity +0.00 emotion 77 %   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #12

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration around average — 0.48, lower than 52 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.42.

At the same time Infatuation goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.48 (lower than 52 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.20 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.75 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.80 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.75, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 23 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.740 before conversion and 0.744 after — it rose by 0.004. Neighbour-to-neighbour the worst pair went 0.833 → 0.749. (The earlier render, with segment 1 left raw, scores 0.594 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.416 in the original and +0.321 after conversion — 77 % of the delta retained, which is most of it. On the other named axis, Infatuation, -0.404 became -0.404.

Quality. Mean predicted overall quality across the segments went 2.82 → 3.03 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.740 → 0.744 +0.004identity cos neighbours 0.833 → 0.749d_b rescored +0.416 → +0.321d_a rescored -0.404 → -0.404d_a mined -0.408d_b mined 0.416min_cos_consec (site) 0.7969min_cos_anchor (site) 0.7461dataset emolialang enspeaker EN_8SWTN1s1q5ktotal 22.7schain gain +0.9 dBseam step 1.8 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, steady, almost no disfluency
(normal-paced, average clarity, formal) Followed in dot notation by the Java bean, name, and then the property.
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.6/10; 4.4s, EN.
EN_8SWTN1s1q5k_W000025 · in -10.9 dBFS · gain -9.1 dB · emolia-01739
(normal-paced, clear, formal, narration) So with this example, we are getting the same hourly rate from the employee Java bean, which is stored as a request attribute.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.3/10; 7.4s, EN.
EN_8SWTN1s1q5k_W000026 · in -10.9 dBFS · gain -9.1 dB · emolia-01739
(measured, clear, formal, monologue) Now there are four possibilities for the scope. These are all accessed by accessing some of the implicit objects that are available to us with our JSP and the expression language.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, quiet background; genuineness 0.5/6; vocal-burst blend 0.1/10; 11.4s, EN.
EN_8SWTN1s1q5k_W000027 · in -13.1 dBFS · gain -6.9 dB · emolia-01739
Concentration ↓  /  Doubtidentity +0.06 emotion REVERSED   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #13

This chain comes from the proxy rule: the same two-sided test as above, but because Doubt is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Doubt barely there — 0.22, lower than 78 % of clips in this corpus — and ends with it strongly present at 0.76, higher than 76 % of clips in this corpus. That is a total rise of 0.54.

At the same time Concentration goes the other way, from 0.93 (higher than 92 % of clips in this corpus) to 0.49 (right about the corpus median), a change of -0.43. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.15, then +0.19, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.80 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 26 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.718 before conversion and 0.780 after — it rose by 0.062. Neighbour-to-neighbour the worst pair went 0.752 → 0.836. (The earlier render, with segment 1 left raw, scores 0.726 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Doubt moved +0.541 in the original and -0.205 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Concentration, -0.433 became -0.260.

Quality. Mean predicted overall quality across the segments went 3.06 → 3.16 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.718 → 0.780 +0.062identity cos neighbours 0.752 → 0.836d_b rescored +0.541 → -0.205d_a rescored -0.433 → -0.260d_a mined -0.432d_b mined 0.541min_cos_consec (site) 0.7964min_cos_anchor (site) 0.8345dataset emolialang zhspeaker ZH_B00046_S04772total 24.7schain gain +1.1 dBseam step 0.7 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady
(concentration · measured, some disfluency, clear, didactic) 请求叶知县把雷鸣陈亮从牢里放出来,派他们到临安去,请济公。叶志宪答应了雷鸣和陈亮,立刻去请。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: didactic, monologue; average recording, no background noise; genuineness 2.1/6; vocal-burst blend 2.6/10; 10.5s, ZH.
ZH_B00046_S04772_W000044 · in -14.7 dBFS · gain -5.3 dB · emolia-03734
(confusion, intoxication altered states of consciousness, emotional numbness · measured, some disfluency, average clarity, didactic) 这时,杨明、赵斌、柳瑞三个人混进了刘香庙的匪巢小西天。
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, intoxication altered states of consciousness, emotional numbness; style: didactic, authoritative; average recording, no background noise; genuineness 2.1/6; vocal-burst blend 3.4/10; 6.6s, ZH.
ZH_B00046_S04772_W000045 · in -14.9 dBFS · gain -5.1 dB · emolia-03734
(normal-paced, no disfluency, clear, authoritative) 刺探秘密的时候,被刘香庙的同党捉住了。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 2.4/10; 3.3s, ZH.
ZH_B00046_S04772_W000046 · in -14.2 dBFS · gain -5.8 dB · emolia-03734
(fast, little disfluency, clear, authoritative) 济公到了玉山,就亲自带领官兵进攻小西天的匪巢。
full caption & clip details
An adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; average recording, no background noise; genuineness 2.2/6; vocal-burst blend 2.9/10; 4.9s, ZH.
ZH_B00046_S04772_W000047 · in -14.0 dBFS · gain -6.0 dB · emolia-03734
Contempt ↓  /  Prideidentity −0.02 emotion 115 %   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #14

This chain comes from the proxy rule: the same two-sided test as above, but because Pride is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Pride around average — 0.53, higher than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.44.

At the same time Contempt goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.50 (right about the corpus median), a change of -0.46. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.13, then -0.02, then +0.11 — not a clean run: step 3 moves back the other way by 0.02 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 63 s · fr · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.879 before conversion and 0.858 after — it fell by 0.022. Neighbour-to-neighbour the worst pair went 0.836 → 0.845. (The earlier render, with segment 1 left raw, scores 0.777 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.724 in the original and +0.830 after conversion — 115 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contempt, -0.456 became -0.344.

Quality. Mean predicted overall quality across the segments went 2.93 → 3.13 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.879 → 0.858 -0.022identity cos neighbours 0.836 → 0.845d_b rescored +0.724 → +0.830d_a rescored -0.456 → -0.344d_a mined -0.456d_b mined 0.438min_cos_consec (site) 0.9417min_cos_anchor (site) 0.9191dataset emolialang frspeaker FR_3a_QqReFVK0total 61.4schain gain +0.1 dBseam step 0.9 dBcrossfades 150/100/100/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, slightly dark, slightly rough, quiet background, normally alert, somewhat unclear
(contempt, astonishment surprise, distress · measured, neutral tension, moderately variable, casual) Ne vous lancez jamais sur un compte réel, surtout sans avoir de formation. (low mumble) Euh, vous avez toutes les chances de perdre de l'argent. Donc, commencez d'abord sur un compte virtuel. Vous ne prenez aucun risque et,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, thin; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, astonishment surprise, distress; style: casual, monologue; below-average recording, quiet background; genuineness 4.4/6; vocal-burst blend 6.0/10; 11.9s, FR.
FR_3a_QqReFVK0_W000008 · in -14.8 dBFS · gain -5.2 dB · emolia-02837
(normal-paced, neutral tension, moderately variable, casual) suivez donc une formation, commencez sur un compte virtuel, et une fois que vous aurez des bons résultats, vous commencez, vous pourrez commencer à gagner de l'argent réellement avec un petit capital, vous pourrez commencer avec 500 euros, 1000 euros, 5000 euros, etc.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: casual, monologue; below-average recording, quiet background; genuineness 4.0/6; vocal-burst blend 8.3/10; 13.4s, FR.
FR_3a_QqReFVK0_W000009 · in -14.9 dBFS · gain -5.1 dB · emolia-02837
(fear, distress · measured, slightly relaxed, fairly steady, casual) Et je peux vous assurer que quand vous aurez des bons résultats, vous n'aurez aucun problème pour trouver des capitaux, des gens qui (low mumble) vous font confiance et qui veulent profiter de votre savoir-faire. Donc,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as fear, distress; style: casual, monologue; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 3.5/10; 10.8s, FR.
FR_3a_QqReFVK0_W000010 · in -14.3 dBFS · gain -5.7 dB · emolia-02837
(measured, slightly relaxed, fairly steady, monologue) euh, (low mumble) tout vous est ouvert, toutes les portes vous seront ouvertes une fois que vous serez devenu bon en trading. Et, euh, (low mumble) c'est vraiment le, le métier idéal. Si vous cherchez une activité rémunératrice, une activité bien payée, vous n'aurez pas besoin de chercher un job puisque le job vous l'aurez déjà.
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 4.5/10; 18.0s, FR.
FR_3a_QqReFVK0_W000011 · in -15.7 dBFS · gain -4.3 dB · emolia-02837
(pride, triumph · measured, slightly relaxed, fairly steady, casual) Voilà, j'espère que vous avez aimé cette vidéo. Likez-la si c'est le cas. Partagez-la. Ça m'encouragera à faire de nouvelles vidéos.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as pride, triumph; style: casual, playful; below-average recording, quiet background; genuineness 3.2/6; vocal-burst blend 3.5/10; 8.0s, FR.
FR_3a_QqReFVK0_W000012 · in -13.4 dBFS · gain -6.6 dB · emolia-02837
Interest ↓  /  Intoxication Altered States of Consciousnessidentity +0.16 emotion REVERSED   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #15

This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Intoxication Altered States of Consciousness around average — 0.43, lower than 57 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.43.

At the same time Interest goes the other way, from 0.79 (higher than 79 % of clips in this corpus) to 0.38 (lower than 62 % of clips in this corpus), a change of -0.42. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.18, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.32 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.32 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.32, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 14 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.315 before conversion and 0.478 after — it rose by 0.163. Neighbour-to-neighbour the worst pair went 0.315 → 0.478. (The earlier render, with segment 1 left raw, scores 0.485 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

The emotional move did not survive. Re-scored end to end, Intoxication Altered States of Consciousness moved +0.430 in the original and -0.375 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Interest, -0.418 became -0.304.

Quality. Mean predicted overall quality across the segments went 2.49 → 2.74 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.315 → 0.478 +0.163identity cos neighbours 0.315 → 0.478d_b rescored +0.430 → -0.375d_a rescored -0.418 → -0.304d_a mined -0.418d_b mined 0.430min_cos_consec (site) 0.3176min_cos_anchor (site) 0.3176dataset emolialang enspeaker EN_oXsFNl-N834total 13.4schain gain +3.2 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(slightly relaxed, some disfluency, average clarity, authoritative) Yep, we've got one more question and one more comment. So the question's from Lizzie Smith, who you know.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, formal; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 1.0/10; 4.6s, EN.
EN_oXsFNl-N834_W000328 · in -17.0 dBFS · gain -3.0 dB · emolia-02280
(doubt · relaxed, frequent disfluency, somewhat unclear, casual) I think, I mean the NDIS funding, I can see it working in two
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as doubt; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.6/10; 4.3s, EN.
EN_oXsFNl-N834_W000330 · in -18.6 dBFS · gain -1.4 dB · emolia-02280
(slightly relaxed, some disfluency, average clarity, casual) In two directions, I mean there's many more than that, but the two directions that come to mind now was...
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.3/10; 4.9s, EN.
EN_oXsFNl-N834_W000331 · in -18.0 dBFS · gain -2.0 dB · emolia-02280
Fear ↓  /  Aweidentity +0.61 emotion 97 %   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #16

This chain comes from the proxy rule: the same two-sided test as above, but because Awe is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Awe around average — 0.47, lower than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.52.

At the same time Fear goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.40 (lower than 60 % of clips in this corpus), a change of -0.58. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.00, then +0.00, then +0.52, then -0.01 — not a clean run: step 4 moves back the other way by 0.01 before the chain recovers.

The largest step is 0.52, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 34 s · snippets

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.002 before conversion and 0.609 after — it rose by 0.606. Neighbour-to-neighbour the worst pair went 0.002 → 0.693. (The earlier render, with segment 1 left raw, scores 0.448 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.515 in the original and +0.501 after conversion — 97 % of the delta retained, which is essentially all of it. On the other named axis, Fear, -0.582 became -0.518.

Quality. Mean predicted overall quality across the segments went 2.73 → 2.90 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.002 → 0.609 +0.606identity cos neighbours 0.002 → 0.693d_b rescored +0.515 → +0.501d_a rescored -0.582 → -0.518d_a mined -0.582d_b mined 0.515min_cos_consec (site) —min_cos_anchor (site) —dataset snippetslang undspeaker batch123_part3_batch123_patotal 32.5schain gain +3.9 dBseam step 1.3 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, good recording, no background noise, clear, light breath
(fear, relief, confusion · brisk, normally alert, slightly relaxed, dramatic) Carl McDonald along with other military personnel went through the week-long training process that anyone can do, even today. But the gateway report ended with a final warning that nobody really expected.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as fear, relief, confusion; style: dramatic, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.8/10; 10.6s.
batch123_part3_batch123_part3_chunk_209_1_299837 · in -17.1 dBFS · gain -2.9 dB · snippets-00126
(brisk, normally alert, slightly relaxed, casual) that the military should learn to adjust reality for what he calls
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, storytelling; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 4.9/10; 3.5s.
batch123_part3_batch123_part3_chunk_209_1_300006 · in -16.7 dBFS · gain -3.3 dB · snippets-00126
(emotional numbness · brisk, energised, neutral tension, casual) to pursue self-knowledge, and remove any personal prejudices that could be blocking their progress. The ultimate goal is patterning.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; clear, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as emotional numbness; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 2.9/6; vocal-burst blend 4.1/10; 7.0s.
batch123_part3_batch123_part3_chunk_209_1_300076 · in -17.3 dBFS · gain -2.7 dB · snippets-00126
(awe, infatuation · measured, normally alert, slightly relaxed, narration) Human consciousness passes through the looking glass of time space after the fashion of Alice, beginning her journey into wonderland.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, infatuation; style: narration, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.7/10; 8.0s.
batch123_part3_batch123_part3_chunk_209_1_300170 · in -21.1 dBFS · gain +1.1 dB · snippets-00126
(awe, astonishment surprise · normal-paced, normally alert, slightly relaxed, storytelling) They could travel anywhere in the universe, and at any point in time.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as awe, astonishment surprise; style: storytelling, casual; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 4.5/10; 4.0s.
batch123_part3_batch123_part3_chunk_209_1_300433 · in -18.1 dBFS · gain -1.9 dB · snippets-00126
Contemplation ↓  /  Infatuationidentity −0.01 emotion 6 %   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #17

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation around average — 0.45, lower than 55 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.44.

At the same time Contemplation goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.32 (lower than 68 % of clips in this corpus), a change of -0.54. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.15, then +0.03, then +0.09, then +0.17 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.70 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.72 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.70, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 33 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.657 before conversion and 0.650 after — it fell by 0.007. Neighbour-to-neighbour the worst pair went 0.633 → 0.709. (The earlier render, with segment 1 left raw, scores 0.697 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.438 in the original and +0.025 after conversion — 6 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Contemplation, -0.538 became -0.483.

Quality. Mean predicted overall quality across the segments went 2.92 → 3.04 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.657 → 0.650 -0.007identity cos neighbours 0.633 → 0.709d_b rescored +0.438 → +0.025d_a rescored -0.538 → -0.483d_a mined -0.537d_b mined 0.438min_cos_consec (site) 0.7195min_cos_anchor (site) 0.6997dataset emolialang zhspeaker ZH_B00081_S00416total 31.7schain gain +2.3 dBseam step 2.3 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · fairly smooth, no background noise, normally alert, slightly relaxed
(measured, steady, no disfluency, formal) 国有企业也有有效的阻碍着经济和政治制度的创新。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, no disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; average recording, no background noise; genuineness 1.6/6; vocal-burst blend 1.6/10; 5.5s, ZH.
ZH_B00081_S00416_W000043 · in -19.1 dBFS · gain -0.9 dB · emolia-04088
(measured, steady, frequent disfluency, monologue) 国有企业好比农民的自留地,自给自足,不受外界环境过度的影响。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, narrow pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 1.6/6; vocal-burst blend 1.4/10; 7.4s, ZH.
ZH_B00081_S00416_W000044 · in -19.5 dBFS · gain -0.5 dB · emolia-04088
(measured, steady, no disfluency, narration) 历朝历代,垄断关键的工业和商业政府,所需资源大多来自这个部门。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; average recording, no background noise; genuineness 1.6/6; vocal-burst blend 2.8/10; 6.8s, ZH.
ZH_B00081_S00416_W000045 · in -20.0 dBFS · gain -0.0 dB · emolia-04088
(slow, steady, no disfluency, monologue) 这导致了中国历史上从来没有发展出一个比较有效的财政金融和信用等制度体系。
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is warm, slightly dark, fairly smooth, thin; somewhat unclear, no disfluency, narrow pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, whispered; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.7/10; 9.5s, ZH.
ZH_B00081_S00416_W000046 · in -20.2 dBFS · gain +0.2 dB · emolia-04088
(measured, fairly steady, no disfluency, formal) 在西方,因为政府没有自己的企业。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; very good recording, no background noise; genuineness 1.9/6; vocal-burst blend 3.1/10; 3.3s, ZH.
ZH_B00081_S00416_W000047 · in -19.3 dBFS · gain -0.7 dB · emolia-04088
Contemplation ↓  /  Sexual Lustidentity −0.03 emotion 89 %   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #18

This chain comes from the proxy rule: the same two-sided test as above, but because Sexual Lust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Sexual Lust below average — 0.33, lower than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.60.

At the same time Contemplation goes the other way, from 0.82 (higher than 82 % of clips in this corpus) to 0.40 (lower than 60 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.12, then +0.25 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 27 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.862 before conversion and 0.836 after — it fell by 0.026. Neighbour-to-neighbour the worst pair went 0.895 → 0.823. (The earlier render, with segment 1 left raw, scores 0.775 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sexual Lust moved +0.600 in the original and +0.534 after conversion — 89 % of the delta retained, which is most of it. On the other named axis, Contemplation, -0.411 became -0.553.

Quality. Mean predicted overall quality across the segments went 2.94 → 3.23 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.862 → 0.836 -0.026identity cos neighbours 0.895 → 0.823d_b rescored +0.600 → +0.534d_a rescored -0.411 → -0.553d_a mined -0.411d_b mined 0.600min_cos_consec (site) 0.8932min_cos_anchor (site) 0.8838dataset emolialang zhspeaker ZH_B00080_S08566total 26.3schain gain +2.8 dBseam step 1.7 dBcrossfades 150/100/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · fairly smooth, no background noise, slightly relaxed, no disfluency, clear
(measured, normally alert, steady, whispered) 每一个人必须吃许多不同的维生素、矿物质、蛋白质、脂肪和碳水化合物。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, narrow pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: whispered, ASMR; average recording, no background noise; genuineness 0.8/6; vocal-burst blend 4.0/10; 8.3s, ZH.
ZH_B00080_S08566_W000013 · in -16.6 dBFS · gain -3.4 dB · emolia-04078
(emotional numbness, confusion, doubt · measured, normally alert, fairly steady, monologue) 长期缺乏这些合理的营养成分,体内的一些器官就会出现严重的故障,你就会得病。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, confusion, doubt; style: monologue, formal; average recording, no background noise; genuineness 0.5/6; vocal-burst blend 3.5/10; 8.1s, ZH.
ZH_B00080_S08566_W000014 · in -15.8 dBFS · gain -4.2 dB · emolia-04078
(slow, normally alert, steady, ASMR) 每一种营养需要多少才是合理的呢?这个问题。
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: ASMR, whispered; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 3.9/10; 5.4s, ZH.
ZH_B00080_S08566_W000015 · in -15.6 dBFS · gain -4.4 dB · emolia-04078
(sexual lust, malevolence malice, confusion · slow, very low-energy, fairly steady, narration) 不可能有适用于所有人的答案,这是很清楚的。
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is slightly warm, slightly dark, fairly smooth, very full; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, fairly guarded; reads as sexual lust, malevolence malice, confusion; style: narration, whispered; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 8.0/10; 5.0s, ZH.
ZH_B00080_S08566_W000016 · in -16.7 dBFS · gain -3.3 dB · emolia-04078
Impatience and Irritability ↓  /  Interestidentity −0.03 emotion 94 %   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #19

This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Interest barely there — 0.16, lower than 84 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.83.

At the same time Impatience and Irritability goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.39 (lower than 61 % of clips in this corpus), a change of -0.55. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.22, then +0.18, then +0.18 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 54 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.683 before conversion and 0.649 after — it fell by 0.035. Neighbour-to-neighbour the worst pair went 0.770 → 0.784. (The earlier render, with segment 1 left raw, scores 0.593 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.829 in the original and +0.783 after conversion — 94 % of the delta retained, which is essentially all of it. On the other named axis, Impatience and Irritability, -0.550 became -0.744.

Quality. Mean predicted overall quality across the segments went 2.99 → 3.34 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.683 → 0.649 -0.035identity cos neighbours 0.770 → 0.784d_b rescored +0.829 → +0.783d_a rescored -0.550 → -0.744d_a mined -0.550d_b mined 0.829min_cos_consec (site) 0.8138min_cos_anchor (site) 0.8138dataset emolialang zhspeaker ZH_B00063_S09717total 52.6schain gain +1.7 dBseam step 1.1 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(impatience and irritability, distress · normal-paced, no disfluency, clear, authoritative) 数次帮助慧珠逃跑,最后机智的慧珠关闭了机舱内的灯。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as impatience and irritability, distress; style: authoritative, formal; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 3.5/10; 4.1s, ZH.
ZH_B00063_S09717_W000026 · in -19.6 dBFS · gain -0.4 dB · emolia-03909
(fast, little disfluency, average clarity, authoritative) 混淆其视线,拿到了警察的配枪,射中了罪犯,现在幸存者就剩下慧珠和瞎了眼的机长。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, dramatic; good recording, quiet background; genuineness 1.6/6; vocal-burst blend 3.8/10; 6.3s, ZH.
ZH_B00063_S09717_W000027 · in -19.7 dBFS · gain -0.3 dB · emolia-03909
(brisk, some disfluency, average clarity, didactic) 就当两人以为一切都结束,时命大的罪犯苏醒杀掉了机长,究竟慧珠能否逃脱呢?我想是可以的,毕竟主角光环和女鬼加持的威力也不是盖的。
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: didactic, authoritative; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 5.9/10; 10.6s, ZH.
ZH_B00063_S09717_W000028 · in -19.9 dBFS · gain -0.1 dB · emolia-03909
(infatuation · normal-paced, almost no disfluency, clear, monologue) 第三个故事,绿豆红豆这个故事有点像暗黑版的灰姑娘。灰姑娘公治肤白貌美,清秀可人。但她也有一个贪得无厌的继母和蛮横的妹妹公治,被年近六十岁的财阀会长相中,取来做第六位田房。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation; style: monologue, didactic; average recording, no background noise; genuineness 1.1/6; vocal-burst blend 6.1/10; 15.0s, ZH.
ZH_B00063_S09717_W000029 · in -19.8 dBFS · gain -0.2 dB · emolia-03909
(interest, doubt, concentration · brisk, almost no disfluency, clear, monologue) 别看会长的年事已高,可外表和身材与三十岁出头的年轻男子无异。为了驻颜不老会长,常年会吃一种新鲜的肉类小菜,据说极为名贵公治。傍上了这样一位金主。就连对继母说话也硬气了三分。可与此同时,继母和他那毫无血眼的妹妹也打起了算盘。
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, doubt, concentration; style: monologue, authoritative; good recording, quiet background; genuineness 0.6/6; vocal-burst blend 6.6/10; 17.4s, ZH.
ZH_B00063_S09717_W000030 · in -19.7 dBFS · gain -0.3 dB · emolia-03909
Concentration ↓  /  Infatuationidentity +0.04 emotion 85 %   proxy_taillift__PXR__T0.40__C0.25__INTERNAL · #20

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation below average — 0.33, lower than 67 % of clips in this corpus — and ends with it strongly present at 0.75, higher than 75 % of clips in this corpus. That is a total rise of 0.42.

At the same time Concentration goes the other way, from 0.81 (higher than 81 % of clips in this corpus) to 0.39 (lower than 61 % of clips in this corpus), a change of -0.42. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 18 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.729 before conversion and 0.772 after — it rose by 0.042. Neighbour-to-neighbour the worst pair went 0.788 → 0.772. (The earlier render, with segment 1 left raw, scores 0.650 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.423 in the original and +0.358 after conversion — 85 % of the delta retained, which is most of it. On the other named axis, Concentration, -0.424 became -0.366.

Quality. Mean predicted overall quality across the segments went 2.92 → 3.02 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.729 → 0.772 +0.042identity cos neighbours 0.788 → 0.772d_b rescored +0.423 → +0.358d_a rescored -0.424 → -0.366d_a mined -0.424d_b mined 0.423min_cos_consec (site) 0.8588min_cos_anchor (site) 0.8419dataset emolialang zhspeaker ZH_B00057_S05199total 17.3schain gain +2.2 dBseam step 0.7 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(measured, frequent disfluency, somewhat unclear, didactic) 所以说,人们做的只是去寻找完美的人,而不是努力让自己成为完美的人。
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, no background noise; genuineness 2.6/6; vocal-burst blend 1.7/10; 9.5s, ZH.
ZH_B00057_S05199_W000024 · in -20.0 dBFS · gain -0.0 dB · emolia-03851
(normal-paced, no disfluency, clear, authoritative) 而当你试图改变对方的时候,这种做法只会引起争吵。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 2.0/10; 4.7s, ZH.
ZH_B00057_S05199_W000025 · in -18.4 dBFS · gain -1.6 dB · emolia-03851
(normal-paced, some disfluency, average clarity, authoritative) 而最好的办法呢,是应该先改变自己。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 0.5/10; 3.5s, ZH.
ZH_B00057_S05199_W000026 · in -18.9 dBFS · gain -1.1 dB · emolia-03851