proxy_taillift__PXR__T0.20__C0.25__INTERNAL — voice-corrected

Manifest tier. proxy_taillift, rule PXR, T=0.2, step cap 0.25. Population 1,233,330 chains (12,315 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 966,487.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_proxy_taillift__PXR__T0.20__C0.25__INTERNAL.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
38segments re-voiced
0.803 → 0.848median worst-to-anchor identity cosine
90 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Contemplation ↓  /  Fearidentity −0.01 emotion 52 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #1

This chain comes from the proxy rule: the same two-sided test as above, but because Fear is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fear strongly present — 0.78, higher than 78 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.22.

At the same time Contemplation goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.74 (higher than 74 % of clips in this corpus), a change of -0.24. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.14, then +0.07 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.81 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.81 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 30 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.732 before conversion and 0.721 after — it fell by 0.011. Neighbour-to-neighbour the worst pair went 0.732 → 0.721. (The earlier render, with segment 1 left raw, scores 0.537 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.216 in the original and +0.113 after conversion — 52 % of the delta retained. On the other named axis, Contemplation, -0.237 became -0.295.

Quality. Mean predicted overall quality across the segments went 2.71 → 3.08 (+0.37) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.732 → 0.721 -0.011identity cos neighbours 0.732 → 0.721d_b rescored +0.216 → +0.113d_a rescored -0.237 → -0.295d_a mined -0.237d_b mined 0.216min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_aTvBpcHJBlAtotal 29.8schain gain +3.0 dBseam step 0.5 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · fairly smooth, average recording, quiet background
(contemplation, doubt · normal-paced, normally alert, slightly relaxed, monologue) As much as we think often about energy being kind of this, you know, airy-fairy, psychobabbily kind of thing,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contemplation, doubt; style: monologue, casual; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.9/10; 6.4s, EN.
EN_aTvBpcHJBlA_W000583 · in -15.9 dBFS · gain -4.1 dB · emolia-02628
(sexual lust, thankfulness gratitude, fear · brisk, energised, slightly relaxed, dramatic) If you've ever walked into a room where two people are at odds with one another, you will know that energy is real because you feel it in the air. So to be a teacher on the front lines, I mean to be in schools period right now.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as sexual lust, thankfulness gratitude, fear; style: dramatic, casual; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 3.1/10; 13.5s, EN.
EN_aTvBpcHJBlA_W000584 · in -15.7 dBFS · gain -4.3 dB · emolia-02628
(fear, affection, awe · normal-paced, very low-energy, relaxed, casual) All of this is hanging heavy in the air and you breathe in energetically that air day in and day out. (ahem) Uhm, that's, that's a big thing for us to be talking about.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is mildly positive, submissive, slightly vulnerable; reads as fear, affection, awe; style: casual, monologue; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 4.1/10; 10.2s, EN.
EN_aTvBpcHJBlA_W000585 · in -16.6 dBFS · gain -3.4 dB · emolia-02628
Hope Enthusiasm Optimism ↓  /  Painidentity +0.05 emotion 35 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #2

This chain comes from the proxy rule: the same two-sided test as above, but because Pain is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Pain clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.23.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.24. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.73 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.73 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 17 s · en · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.826 before conversion and 0.874 after — it rose by 0.048. Neighbour-to-neighbour the worst pair went 0.826 → 0.874. (The earlier render, with segment 1 left raw, scores 0.783 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.229 in the original and +0.080 after conversion — 35 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Hope Enthusiasm Optimism, -0.244 became -0.116.

Quality. Mean predicted overall quality across the segments went 2.31 → 2.90 (+0.58) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.826 → 0.874 +0.048identity cos neighbours 0.826 → 0.874d_b rescored +0.229 → +0.080d_a rescored -0.244 → -0.116d_a mined -0.244d_b mined 0.229min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_K-ROB1hCESutotal 16.7schain gain +5.8 dBseam step 0.2 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, some disfluency
(hope enthusiasm optimism, interest · normal-paced, light breath, monologue, casual) We'll explore some questions about civic tech and find the action with the help of some fantastic speakers and with all of you having the chance to share your thoughts as well.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism, interest; style: monologue, casual; good recording, quiet background; genuineness 3.2/6; vocal-burst blend 2.0/10; 7.8s, EN.
EN_K-ROB1hCESu_W000008 · in -19.5 dBFS · gain -0.5 dB · emolia-02292
(pain, concentration, pride · measured, normal breath, monologue, whispered) And then we'll think about what might help solve some of the challenges that we've surfaced with a view to commissioning a solution. Some quick housekeeping first.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as pain, concentration, pride; style: monologue, whispered; average recording, no background noise; genuineness 2.3/6; vocal-burst blend 2.0/10; 9.2s, EN.
EN_K-ROB1hCESu_W000009 · in -20.4 dBFS · gain +0.4 dB · emolia-02292
Fear ↓  /  Concentrationidentity +0.08 emotion 59 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #3

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.33.

At the same time Fear goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.63 (higher than 63 % of clips in this corpus), a change of -0.23. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.14, then +0.19 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.78 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.78 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.78, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 37 s · sv · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.797 before conversion and 0.878 after — it rose by 0.080. Neighbour-to-neighbour the worst pair went 0.797 → 0.832. (The earlier render, with segment 1 left raw, scores 0.747 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.320 in the original and +0.189 after conversion — 59 % of the delta retained. On the other named axis, Fear, -0.237 became -0.649.

Quality. Mean predicted overall quality across the segments went 3.06 → 3.22 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.797 → 0.878 +0.080identity cos neighbours 0.797 → 0.832d_b rescored +0.320 → +0.189d_a rescored -0.237 → -0.649d_a mined -0.228d_b mined 0.326min_cos_consec (site) 0.7809min_cos_anchor (site) 0.7809dataset podcastlang svspeaker 465629total 36.1schain gain +1.8 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · slightly dark, fairly smooth, average recording, quiet background, slightly relaxed, fairly steady, frequent disfluency, fairly narrow pitch
(normal-paced, normally alert, average clarity, whispered) ingen risk ökning eller snarare minskad risk i randomiserade studier. Om man lägger till gulkapshormon.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, fairly smooth, thin; average clarity, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: whispered, monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.9/10; 9.2s, SV.
465629_00132912 · in -26.8 dBFS · gain +6.8 dB · podcast-05655
(distress, helplessness · measured, very low-energy, average clarity, whispered) Kombinationsbehandling så har man sett möjligen lite ökad (low mumble) risk för bröskkan.
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, fairly smooth, thin; average clarity, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; reads as distress, helplessness; style: whispered, ASMR; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.2/10; 8.3s, SV.
465629_00133968 · in -25.9 dBFS · gain +5.9 dB · podcast-05651
(concentration, intoxication altered states of consciousness · normal-paced, normally alert, somewhat unclear, whispered) (ahem) Sen har också medosen med hur länge man använder. Det finns andra riskfaktorer, till exempel de som har biemin över 30 så risken ökar. Det finns andra faktorer också som har med det och med riskökningen att göra.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, neutral openness; reads as concentration, intoxication altered states of consciousness; style: whispered, casual; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 2.4/10; 18.9s, SV.
465629_00135468 · in -27.6 dBFS · gain +7.7 dB · podcast-05654
Emotional Numbness ↓  /  Interestidentity −0.10 emotion 93 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #4

This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Interest clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.30.

At the same time Emotional Numbness goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.14, then +0.09, then +0.07 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 46 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.954 before conversion and 0.853 after — it fell by 0.102. Neighbour-to-neighbour the worst pair went 0.943 → 0.853. (The earlier render, with segment 1 left raw, scores 0.828 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.304 in the original and +0.284 after conversion — 93 % of the delta retained, which is essentially all of it. On the other named axis, Emotional Numbness, -0.251 became -0.342.

Quality. Mean predicted overall quality across the segments went 3.08 → 3.24 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.954 → 0.853 -0.102identity cos neighbours 0.943 → 0.853d_b rescored +0.304 → +0.284d_a rescored -0.251 → -0.342d_a mined -0.252d_b mined 0.304min_cos_consec (site) 0.9446min_cos_anchor (site) 0.9574dataset emolialang enspeaker EN_36ddh9kCW3ototal 44.9schain gain +2.9 dBseam step 0.6 dBcrossfades 100/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, good recording, normally alert, slightly relaxed, fairly steady, almost no disfluency, clear
(emotional numbness · normal-paced, moderate pitch range, light breath, narration) Exactly the same happens on Earth, only without the happy boxes. Earthlings called artists have learned to imitate the stimuli of beauty in paintings, sculptures, and so on.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: narration, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.7/10; 9.6s, EN.
EN_36ddh9kCW3o_W000061 · in -15.9 dBFS · gain -4.1 dB · emolia-01942
(sourness · normal-paced, moderate pitch range, minimal breath, formal) These objects are called artworks. Art is related to natural beauty like diet drinks are related to actual sugar, or porn to actual sex. They imitate the stimuli of pleasure, but out of context and deprived of their utility.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as sourness; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.9/10; 12.4s, EN.
EN_36ddh9kCW3o_W000062 · in -16.3 dBFS · gain -3.7 dB · emolia-01942
(awe · measured, moderate pitch range, light breath, formal) Not surprisingly, earthling art often focuses on beautiful people, and on the abundance of nature, landscapes, flowers, fruits, and all kinds of living creatures. Except for, the microbes.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe; style: formal, newsreading; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.7/10; 11.6s, EN.
EN_36ddh9kCW3o_W000063 · in -17.3 dBFS · gain -2.7 dB · emolia-01942
(interest · normal-paced, wide pitch range, light breath, authoritative) If you are interested in art, visit an art museum. However, you should avoid modern art, as it often challenges or distorts the concept of beauty. Better start with classical art, it's much easier to understand.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, full; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as interest; style: authoritative, dramatic; good recording, quiet background; genuineness 0.3/6; vocal-burst blend 4.4/10; 11.9s, EN.
EN_36ddh9kCW3o_W000066 · in -17.2 dBFS · gain -2.8 dB · emolia-01942
Contemplation ↓  /  Emotional Numbnessidentity −0.06 emotion 41 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #5

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness clearly present — 0.73, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.23.

At the same time Contemplation goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.10 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.87 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.87 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 46 s · de · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.809 before conversion and 0.749 after — it fell by 0.061. Neighbour-to-neighbour the worst pair went 0.797 → 0.729. (The earlier render, with segment 1 left raw, scores 0.537 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.230 in the original and +0.094 after conversion — 41 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Contemplation, -0.301 became -0.152.

Quality. Mean predicted overall quality across the segments went 2.82 → 3.28 (+0.46) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.809 → 0.749 -0.061identity cos neighbours 0.797 → 0.729d_b rescored +0.230 → +0.094d_a rescored -0.301 → -0.152d_a mined -0.300d_b mined 0.229min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_XH7jFb1N1Zktotal 45.6schain gain +0.7 dBseam step 0.8 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, balanced body, average recording, quiet background, measured, slightly relaxed, somewhat unclear
(contemplation, concentration, disappointment · subdued, steady, frequent disfluency, whispered) Und der Funktion zu Hashimoto, et cetera. Die Schulmedizin meint nun, das könnte man mit der Gabe von Schilddrüsenhormonen ausgleichen. Doch hier sagt Anthony William White, gefehlt, die Hormone wirken wie Steroidhormone, welche das Immunsystem rapide.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, concentration, disappointment; style: whispered, monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 1.6/10; 19.0s, DE.
DE_XH7jFb1N1Zk_W000047 · in -17.5 dBFS · gain -2.5 dB · emolia-00215
(concentration, interest, fear · subdued, fairly steady, some disfluency, monologue) verlangsamen. Wenn normalerweise genügend Schilddrüsenhormone produziert werden, geht ein Signal an bestimmte Zellen, sodass diese Glukose aufnehmen sollten, um eben Energie für die Reparatur und Vermehrung zu gewinnen. Werden jedoch zu wenig Hormone produziert, gilt das
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, interest, fear; style: monologue, didactic; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.0/10; 19.4s, DE.
DE_XH7jFb1N1Zk_W000048 · in -18.8 dBFS · gain -1.2 dB · emolia-00215
(emotional numbness, concentration · normally alert, fairly steady, frequent disfluency, didactic) an die Zellen, so geht's dann das Signal an die Zellen, (low mumble) die Energieversorgung quasi genügt.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, concentration; style: didactic, monologue; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 0.0/10; 7.6s, DE.
DE_XH7jFb1N1Zk_W000049 · in -15.9 dBFS · gain -4.1 dB · emolia-00215
Contempt ↓  /  Shameidentity +0.66 emotion 93 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #6

This chain comes from the proxy rule: the same two-sided test as above, but because Shame is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Shame clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.25.

At the same time Contempt goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.78 (higher than 78 % of clips in this corpus), a change of -0.22. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.11, then +0.15 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 52 s · pt · eurospeech

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.206 before conversion and 0.869 after — it rose by 0.663. Neighbour-to-neighbour the worst pair went 0.249 → 0.802. (The earlier render, with segment 1 left raw, scores 0.504 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.254 in the original and +0.237 after conversion — 93 % of the delta retained, which is essentially all of it. On the other named axis, Contempt, -0.215 became -0.219.

Quality. Mean predicted overall quality across the segments went 2.89 → 3.13 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.206 → 0.869 +0.663identity cos neighbours 0.249 → 0.802d_b rescored +0.254 → +0.237d_a rescored -0.215 → -0.219d_a mined -0.215d_b mined 0.254min_cos_consec (site) —min_cos_anchor (site) —dataset eurospeechlang ptspeaker portugal_12_2_51total 51.0schain gain +2.8 dBseam step 0.9 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · slightly cool, neutral-bright, thin
(contempt, disgust, bitterness · fast, highly aroused, tense, dramatic) falando com um único concorrente, colocando o Estado nas mãos desse concorrente, vendendo limpo de passivo e por um preço ao desbarato os ativos dos Estaleiros. Não disseram uma palavra sobre isso!
full caption & clip details
A young adult masculine voice; delivery is highly aroused, fast, tense, volatile; timbre is slightly cool, neutral-bright, rough, thin; clear, almost no disfluency, very wide pitch range, audible breath; affect is elated, very dominant, guarded; reads as contempt, disgust, bitterness; style: dramatic, ranting; poor recording, noisy background; genuineness 2.3/6; vocal-burst blend 5.3/10; 14.4s, PT.
portugal_12_2_51_6164640_6178992 · in -14.9 dBFS · gain -5.1 dB · eurospeech-02492
(anger, contempt, malevolence malice · fast, highly aroused, tense, authoritative) sobre isso! Estão interessados nessa via destrutiva, por isso posso concluir este debate dizendo que se o PS começou o enterro, o PSD e o CDS querem ser definitivamente os coveiros dos Estaleiros Navais de Viana do Castelo.
full caption & clip details
An adult masculine voice; delivery is highly aroused, fast, tense, volatile; timbre is slightly cool, neutral-bright, rough, thin; very clear, almost no disfluency, very wide pitch range, audible breath; affect is elated, very dominant, guarded; reads as anger, contempt, malevolence malice; style: authoritative, dramatic; poor recording, noisy background; genuineness 1.6/6; vocal-burst blend 4.4/10; 18.0s, PT.
portugal_12_2_51_6178992_6196951 · in -16.0 dBFS · gain -4.0 dB · eurospeech-02492
(shame, thankfulness gratitude · normal-paced, normally alert, neutral tension, cartoonish) Muito obrigada,
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as shame, thankfulness gratitude; style: cartoonish, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 4.6/10; 19.2s, PT.
portugal_12_2_51_6257232_6276391 · in -16.1 dBFS · gain -3.9 dB · eurospeech-02492
Triumph ↓  /  Contentmentidentity −0.04 emotion 73 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #7

This chain comes from the proxy rule: the same two-sided test as above, but because Contentment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contentment clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.21.

At the same time Triumph goes the other way, from 0.92 (higher than 92 % of clips in this corpus) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.22. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.21 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 40 s · de · podcast

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.894 before conversion and 0.858 after — it fell by 0.036. Neighbour-to-neighbour the worst pair went 0.894 → 0.858. (The earlier render, with segment 1 left raw, scores 0.669 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.216 in the original and +0.157 after conversion — 73 % of the delta retained, which is most of it. On the other named axis, Triumph, -0.215 became -0.071.

Quality. Mean predicted overall quality across the segments went 3.01 → 3.16 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.894 → 0.858 -0.036identity cos neighbours 0.894 → 0.858d_b rescored +0.216 → +0.157d_a rescored -0.215 → -0.071d_a mined -0.215d_b mined 0.213min_cos_consec (site) 0.8899min_cos_anchor (site) 0.8899dataset podcastlang despeaker 869765total 39.7schain gain +1.3 dBseam step 3.8 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(triumph, jealousy and envy · monologue, casual) ein Hörerin schrieb auch, Auf ihrer ersten Party wurde ihr ungefragt in den Schritt gefasst. (low mumble) Oder es wurde sich auch an ihrer Hüfte dann gerieben, irgendwie in Clubs und so. Ja, es geht wirklich schon in der Schule los. Also (low mumble) in der Grundschule liefen der Hörerin die Jungs hinterher, um ihr an die Brüste zu greifen und haben Titten gucken gerufen.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as triumph, jealousy and envy; style: monologue, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 0.0/10; 20.9s, DE.
869765_00274536 · in -20.1 dBFS · gain +0.1 dB · podcast-01768
(contentment, awe, infatuation · monologue, casual) Ja, eine andere schreibt, dass es einfach in der Schule kneipen war, Clubs in der Stadt auf dem Dorf tagsüber oder nachts, dass es einfach überall und viel zu oft passiert. Und dass eine andere schreibt, dass es seit 13.14 ist ganz normal quasi ist, dass es von unangenehm bis gefährlich sich anfühlt.
full caption & clip details
A young adult somewhat feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contentment, awe, infatuation; style: monologue, casual; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 0.0/10; 19.0s, DE.
869765_00276624 · in -21.2 dBFS · gain +1.2 dB · podcast-01760
Concentration ↓  /  Emotional Numbnessidentity +0.00 emotion 50 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #8

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.24.

At the same time Concentration goes the other way, from 0.92 (higher than 92 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.22. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.01 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 30 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.928 before conversion and 0.928 after — it rose by 0.001. Neighbour-to-neighbour the worst pair went 0.882 → 0.890. (The earlier render, with segment 1 left raw, scores 0.771 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.244 in the original and +0.123 after conversion — 50 % of the delta retained. On the other named axis, Concentration, -0.223 became -0.224.

Quality. Mean predicted overall quality across the segments went 3.23 → 3.32 (+0.09) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.928 → 0.928 +0.001identity cos neighbours 0.882 → 0.890d_b rescored +0.244 → +0.123d_a rescored -0.223 → -0.224d_a mined -0.223d_b mined 0.244min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00016_S04318total 29.6schain gain +0.1 dBseam step 0.7 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, slightly dark, fairly smooth, balanced body, no background noise, measured, normally alert, slightly relaxed
(concentration · fairly steady, some disfluency, somewhat unclear, monologue) 是无限期的等待,还是继续打电话呢?人们常说,坚持就有回报,但销售员应该明白。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, whispered; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 3.8/10; 9.6s, ZH.
ZH_B00016_S04318_W000033 · in -19.8 dBFS · gain -0.2 dB · emolia-03440
(steady, no disfluency, somewhat unclear, monologue) 现在的销售拜访约见规则已经改变了。换句话说,职业化的销售流言已经不够了,因为每个人听起来都很职业化。毫无疑问。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; average recording, no background noise; genuineness 1.8/6; vocal-burst blend 2.8/10; 12.1s, ZH.
ZH_B00016_S04318_W000034 · in -19.0 dBFS · gain -1.0 dB · emolia-03440
(steady, no disfluency, clear, monologue) 如果你跟别人一样,那你就只能得到平均收益。我们知道这种结果令人难以接受。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 4.0/10; 8.3s, ZH.
ZH_B00016_S04318_W000035 · in -18.3 dBFS · gain -1.7 dB · emolia-03440
Astonishment Surprise ↓  /  Longingidentity +0.05 emotion 125 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #9

This chain comes from the proxy rule: the same two-sided test as above, but because Longing is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Longing clearly present — 0.58, higher than 58 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.34.

At the same time Astonishment Surprise goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.45. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.14, then +0.07, then +0.14 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 27 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.821 before conversion and 0.869 after — it rose by 0.048. Neighbour-to-neighbour the worst pair went 0.794 → 0.819. (The earlier render, with segment 1 left raw, scores 0.827 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.344 in the original and +0.430 after conversion — 125 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Astonishment Surprise, -0.449 became -0.183.

Quality. Mean predicted overall quality across the segments went 2.95 → 3.08 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.821 → 0.869 +0.048identity cos neighbours 0.794 → 0.819d_b rescored +0.344 → +0.430d_a rescored -0.449 → -0.183d_a mined -0.449d_b mined 0.344min_cos_consec (site) 0.8152min_cos_anchor (site) 0.8296dataset emolialang zhspeaker ZH_B00079_S01115total 25.9schain gain +1.4 dBseam step 1.4 dBcrossfades 150/150/100 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, measured, normally alert, slightly relaxed, fairly steady
(astonishment surprise · some disfluency, average clarity, moderate pitch range, conversational) 先开始是迫于这个李浩的拳脚威胁,你知道吧?不敢不听从他的安排啊。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as astonishment surprise; style: conversational, storytelling; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 4.7/10; 6.5s, ZH.
ZH_B00079_S01115_W000072 · in -15.9 dBFS · gain -4.1 dB · emolia-04067
(some disfluency, average clarity, moderate pitch range, formal) 到后来居然发展到什么呀?心甘情愿的接受这个李浩的,为什么呀?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 4.2/10; 5.0s, ZH.
ZH_B00079_S01115_W000073 · in -16.2 dBFS · gain -3.8 dB · emolia-04067
(some disfluency, somewhat unclear, moderate pitch range, conversational) (low mumble) 这个心态就发生变化了,甚至还会为了这个李浩。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: conversational, casual; average recording, no background noise; genuineness 3.2/6; vocal-burst blend 4.8/10; 3.4s, ZH.
ZH_B00079_S01115_W000074 · in -16.1 dBFS · gain -3.9 dB · emolia-04067
(longing · no disfluency, clear, fairly narrow pitch, monologue) 二零一零年的下半年啊,两个女孩因为对李浩的争风吃醋就动起手来了,李浩呢就帮助这个其中的一名女孩啊将另一个女孩给杀害了。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing; style: monologue, narration; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 3.2/10; 11.5s, ZH.
ZH_B00079_S01115_W000075 · in -15.8 dBFS · gain -4.2 dB · emolia-04067
Triumph ↓  /  Emotional Numbnessidentity −0.12 emotion 141 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #10

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness strongly present — 0.76, higher than 76 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.20.

At the same time Triumph goes the other way, from 0.81 (higher than 81 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.21. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are -0.04, then +0.24 — not a clean run: step 1 moves back the other way by 0.04 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.93 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.93 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 18 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.912 before conversion and 0.787 after — it fell by 0.125. Neighbour-to-neighbour the worst pair went 0.912 → 0.787. (The earlier render, with segment 1 left raw, scores 0.459 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.202 in the original and +0.284 after conversion — 141 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Triumph, -0.205 became -0.225.

Quality. Mean predicted overall quality across the segments went 2.89 → 2.99 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.912 → 0.787 -0.125identity cos neighbours 0.912 → 0.787d_b rescored +0.202 → +0.284d_a rescored -0.205 → -0.225d_a mined -0.205d_b mined 0.202min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_Vej5VSS8RRctotal 17.3schain gain +1.1 dBseam step 0.4 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fairly steady, formal, authoritative) Hubeler Ross completed her training in psychiatry in 1963, and moved to Chicago in 1965
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.6/10; 6.9s, EN.
EN_Vej5VSS8RRc_W000019 · in -15.3 dBFS · gain -4.7 dB · emolia-02243
(fairly steady, formal, authoritative) She became an instructor at the University of Chicago's Pritzker School of Medicine
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; mildly explicit content; genuineness 0.4/6; vocal-burst blend 1.1/10; 4.5s, EN.
EN_Vej5VSS8RRc_W000020 · in -14.4 dBFS · gain -5.6 dB · emolia-02243
(emotional numbness, fear, infatuation · steady, authoritative, formal) She developed there a series of seminars using interviews with terminal patients, which drew
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, fear, infatuation; style: authoritative, formal; good recording, no background noise; mildly explicit content; genuineness 0.4/6; vocal-burst blend 0.0/10; 6.3s, EN.
EN_Vej5VSS8RRc_W000021 · in -14.7 dBFS · gain -5.3 dB · emolia-02243
Fatigue Exhaustion ↓  /  Malevolence Maliceidentity +0.10 emotion 97 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #11

This chain comes from the proxy rule: the same two-sided test as above, but because Malevolence Malice is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Malevolence Malice clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.30.

At the same time Fatigue Exhaustion goes the other way, from 0.83 (higher than 83 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.12, then +0.18 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 23 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.732 before conversion and 0.831 after — it rose by 0.099. Neighbour-to-neighbour the worst pair went 0.732 → 0.812. (The earlier render, with segment 1 left raw, scores 0.755 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Malevolence Malice moved +0.300 in the original and +0.291 after conversion — 97 % of the delta retained, which is essentially all of it. On the other named axis, Fatigue Exhaustion, -0.307 became -0.431.

Quality. Mean predicted overall quality across the segments went 2.96 → 3.04 (+0.08) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.732 → 0.831 +0.099identity cos neighbours 0.732 → 0.812d_b rescored +0.300 → +0.291d_a rescored -0.307 → -0.431d_a mined -0.307d_b mined 0.300min_cos_consec (site) 0.8307min_cos_anchor (site) 0.8725dataset emolialang zhspeaker ZH_B00070_S07348total 22.3schain gain +2.5 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, clear
(fast, moderately variable, almost no disfluency, didactic) 就是抢行人的财物,那也是国法不容啊,可都是掉脑袋的罪,更何况这次拦的又是镖局保的国宝啊。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: didactic, authoritative; average recording, no background noise; genuineness 2.3/6; vocal-burst blend 4.5/10; 11.9s, ZH.
ZH_B00070_S07348_W000005 · in -17.6 dBFS · gain -2.4 dB · emolia-03980
(impatience and irritability, pain, contempt · normal-paced, moderately variable, no disfluency, storytelling) 我们镖局的人呢可是最讲义气。
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, pain, contempt; style: storytelling, dramatic; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 3.1/10; 3.1s, ZH.
ZH_B00070_S07348_W000006 · in -15.7 dBFS · gain -4.3 dB · emolia-03980
(malevolence malice, disgust, impatience and irritability · measured, fairly steady, some disfluency, authoritative) 按理说这件事儿就应该报官,可是我们却没有,因为我们打算私了。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as malevolence malice, disgust, impatience and irritability; style: authoritative, didactic; average recording, no background noise; genuineness 2.7/6; vocal-burst blend 3.7/10; 7.7s, ZH.
ZH_B00070_S07348_W000007 · in -18.1 dBFS · gain -1.9 dB · emolia-03980
Malevolence Malice ↓  /  Painidentity −0.01 emotion 33 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #12

This chain comes from the proxy rule: the same two-sided test as above, but because Pain is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Pain clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.24.

At the same time Malevolence Malice goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.00, then +0.24 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 43 s · lt · eurospeech

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.877 before conversion and 0.871 after — it fell by 0.006. Neighbour-to-neighbour the worst pair went 0.854 → 0.823. (The earlier render, with segment 1 left raw, scores 0.696 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.242 in the original and +0.080 after conversion — 33 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Malevolence Malice, -0.248 became -0.571.

Quality. Mean predicted overall quality across the segments went 2.35 → 3.19 (+0.84) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.877 → 0.871 -0.006identity cos neighbours 0.854 → 0.823d_b rescored +0.242 → +0.080d_a rescored -0.248 → -0.571d_a mined -0.248d_b mined 0.242min_cos_consec (site) —min_cos_anchor (site) —dataset eurospeechlang ltspeaker lithuania_lithuania_14_300total 42.2schain gain +3.0 dBseam step 2.8 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · slightly dark, balanced body, quiet background, measured, fairly narrow pitch
(malevolence malice · subdued, slightly relaxed, fairly steady, monologue) Bet jeigu įstatymas pagerėjo, nes buvo geresnis modelis, mums patarė, kaip pagerinti, tai prieš ką mes šiaušiamės, ant kokių sienų mes lipame ir su kokia demagogija?
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is warm, slightly dark, rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as malevolence malice; style: monologue, storytelling; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.3/10; 15.2s, LT.
lithuania_lithuania_14_30032000_11997279_12012496 · in -12.8 dBFS · gain -7.2 dB · eurospeech-01872
(contempt, bitterness · normally alert, neutral tension, fairly steady, monologue) kaip Lietuvos kaimas išgyvens. Aišku, tą kaimą gali gąsdinti ir jis visas dreba, kad jis neišgyvens, bet (ahem) pagalvokime, jeigu mes rimtai galvojame apie kaimo ateitį, apie Lietuvos ateitį, apie kokį kaimą kalbam.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, bitterness; style: monologue, casual; below-average recording, quiet background; genuineness 4.1/6; vocal-burst blend 3.5/10; 15.6s, LT.
lithuania_lithuania_14_30032000_12024512_12040080 · in -14.2 dBFS · gain -5.8 dB · eurospeech-01872
(pain, concentration · normally alert, slightly relaxed, steady, didactic) Ar tas kaimas neturi keistis? Ar jis turi būti toks, koks dabar, šiandien? Ar nereikia jam padėti keistis?
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as pain, concentration; style: didactic, monologue; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.0/10; 11.8s, LT.
lithuania_lithuania_14_30032000_12040080_12051857 · in -14.7 dBFS · gain -5.3 dB · eurospeech-01872
Impatience and Irritability ↓  /  Interestidentity +0.53 emotion 117 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #13

This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Interest clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.22.

At the same time Impatience and Irritability goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.45. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.16, then +0.01, then +0.04 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.32 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.32 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.32, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 76 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.314 before conversion and 0.846 after — it rose by 0.532. Neighbour-to-neighbour the worst pair went 0.314 → 0.847. (The earlier render, with segment 1 left raw, scores 0.471 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.218 in the original and +0.255 after conversion — 117 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Impatience and Irritability, -0.447 became -0.479.

Quality. Mean predicted overall quality across the segments went 2.94 → 3.23 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.314 → 0.846 +0.532identity cos neighbours 0.314 → 0.847d_b rescored +0.218 → +0.255d_a rescored -0.447 → -0.479d_a mined -0.446d_b mined 0.216min_cos_consec (site) 0.3244min_cos_anchor (site) 0.3244dataset podcastlang enspeaker 480670total 74.8schain gain +5.2 dBseam step 2.5 dBcrossfades 100/100/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, some disfluency, average clarity, light breath
(impatience and irritability, sourness, jealousy and envy · normal-paced, energised, neutral tension, casual) Yeah, it's it's not like you're you're gonna be able to just say, (ahem) ah, you know what, I want this made. And especially if they're now you have all these other companies that are trying to get their chips made to sell their drones, and DJI says, oh, well, let's use these compliance, they would have to almost buy these companies, which I mean they have the resources to do it. They bought hostile blood for crying out loud, you know, which (ahem) is I mean, they were on the freaking room. Like (childlike giggle)
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as impatience and irritability, sourness, jealousy and envy; style: casual, conversational; good recording, quiet background; genuineness 4.3/6; vocal-burst blend 5.3/10; 25.1s, EN.
480670_00113080 · in -19.2 dBFS · gain -0.8 dB · podcast-06225
(disappointment, concentration, distress · measured, subdued, slightly relaxed, whispered) and really start to re-evaluate where your next system is going to come from. Because the people that unfortunately are making the rules do not know the ramifications of of what they have for pub what they have done to public.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as disappointment, concentration, distress; style: whispered, casual; good recording, no background noise; genuineness 3.0/6; vocal-burst blend 3.6/10; 14.4s, EN.
480670_00117232 · in -18.6 dBFS · gain -1.4 dB · podcast-01832
(normal-paced, normally alert, slightly relaxed, conversational) And also, you know, type your equipment to your use cases as well. So a lot of times you see like big firepower coming out for a smaller operation, like, oh, well, (ahem) we're gonna take this M300 out and we're gonna, you know, do something small or minute with it. You really don't need that. You don't need the thermal, you don't need this, you don't need that in certain applications, but in that certain applications you need that.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: conversational, casual; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 7.8/10; 20.7s, EN.
480670_00118720 · in -18.8 dBFS · gain -1.2 dB · podcast-06247
(interest, concentration · brisk, normally alert, slightly relaxed, conversational) So weigh your options. What are 90% of your use cases and match your equipment to meet your use case to try to tailor in that budget a little bit because you're not gonna find something on the market that's gonna meet every one of your use cases. And if you're paying, you know, 20, 30%
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as interest, concentration; style: conversational, casual; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 5.9/10; 15.1s, EN.
480670_00120783 · in -18.8 dBFS · gain -1.2 dB · podcast-01835
Astonishment Surprise ↓  /  Disappointmentidentity +0.04 emotion 100 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #14

This chain comes from the proxy rule: the same two-sided test as above, but because Disappointment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Disappointment below average — 0.39, lower than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.59.

At the same time Astonishment Surprise goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.75 (higher than 75 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.59 — a single step, so there is no internal shape to speak of.

The largest step is 0.59, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 27 s · en · podcast

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.701 before conversion and 0.736 after — it rose by 0.035. Neighbour-to-neighbour the worst pair went 0.701 → 0.736. (The earlier render, with segment 1 left raw, scores 0.715 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disappointment moved +0.594 in the original and +0.593 after conversion — 100 % of the delta retained, which is essentially all of it. On the other named axis, Astonishment Surprise, -0.240 became -0.205.

Quality. Mean predicted overall quality across the segments went 2.61 → 2.94 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.701 → 0.736 +0.035identity cos neighbours 0.701 → 0.736d_b rescored +0.594 → +0.593d_a rescored -0.240 → -0.205d_a mined -0.247d_b mined 0.593min_cos_consec (site) 0.8056min_cos_anchor (site) 0.8056dataset podcastlang enspeaker 635312total 26.3schain gain +4.3 dBseam step 1.5 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, moderately variable, some disfluency, average clarity
(astonishment surprise, amusement, hope enthusiasm optimism · brisk, energised, slightly tense, casual) But I got a funny story to tell you about G, real quick. G was crazy, right? We all know this how crazy G was.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, fairly guarded; reads as astonishment surprise, amusement, hope enthusiasm optimism; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.6/6; vocal-burst blend 6.5/10; 5.6s, EN.
635312_00079568 · in -25.3 dBFS · gain +5.3 dB · podcast-05603
(disappointment, triumph, jealousy and envy · normal-paced, normally alert, neutral tension, casual) No, he did not play. He did not play around. So my senior year, I was skipping the first two classes of every single day. Because my parents would go to work. My parents would leave for work at like set at like 6 30 in the morning, right? They'd be out. I lived two blocks away from cross. So I was supposed to walk to school every day. It was my senior year. I was like, fuck that.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as disappointment, triumph, jealousy and envy; style: casual, conversational; average recording, quiet background; genuineness 5.6/6; vocal-burst blend 10.0/10; 20.8s, EN.
635312_00080224 · in -26.0 dBFS · gain +6.0 dB · podcast-05231
Fear ↓  /  Contemplationidentity +0.03 emotion 58 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #15

This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contemplation clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.23.

At the same time Fear goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.21. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.85 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.85 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 21 s · en · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.797 before conversion and 0.825 after — it rose by 0.029. Neighbour-to-neighbour the worst pair went 0.797 → 0.825. (The earlier render, with segment 1 left raw, scores 0.627 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.226 in the original and +0.132 after conversion — 58 % of the delta retained. On the other named axis, Fear, -0.214 became -0.368.

Quality. Mean predicted overall quality across the segments went 2.36 → 3.08 (+0.72) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.797 → 0.825 +0.029identity cos neighbours 0.797 → 0.825d_b rescored +0.226 → +0.132d_a rescored -0.214 → -0.368d_a mined -0.214d_b mined 0.226min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_vKrwQ_58ZgItotal 20.5schain gain +2.8 dBseam step 0.2 dBcrossfades 100 ms
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, slightly dark, fairly smooth, balanced body, average recording, quiet background, slightly relaxed, fairly steady
(fear, concentration · measured, subdued, fairly narrow pitch, monologue) (low mumble) A significant part of the Bitcoin ecosystem have not been paying attention to this DeFi (ahem) movement. I mean (low mumble) Ethereum has been (low mumble) live for several years now. (low mumble)
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as fear, concentration; style: monologue, casual; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 1.9/10; 13.7s, EN.
EN_vKrwQ_58ZgI_W000308 · in -19.0 dBFS · gain -1.0 dB · emolia-02167
(contemplation · normal-paced, normally alert, moderate pitch range, monologue) I would say (low mumble) that the DeFi movement is like the most interesting things we've seen so far.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, neutral openness; reads as contemplation; style: monologue, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.4/10; 7.0s, EN.
EN_vKrwQ_58ZgI_W000309 · in -17.7 dBFS · gain -2.3 dB · emolia-02167
Confusion ↓  /  Infatuationidentity +0.38 emotion 115 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #16

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it strongly present at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.23.

At the same time Confusion goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.75 (higher than 75 % of clips in this corpus), a change of -0.23. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.10, then -0.05, then +0.09, then +0.09 — not a clean run: step 2 moves back the other way by 0.05 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 67 s · da · eurospeech

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.468 before conversion and 0.850 after — it rose by 0.381. Neighbour-to-neighbour the worst pair went 0.483 → 0.816. (The earlier render, with segment 1 left raw, scores 0.714 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.226 in the original and +0.260 after conversion — 115 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Confusion, -0.233 became -0.273.

Quality. Mean predicted overall quality across the segments went 3.02 → 3.38 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.468 → 0.850 +0.381identity cos neighbours 0.483 → 0.816d_b rescored +0.226 → +0.260d_a rescored -0.233 → -0.273d_a mined -0.233d_b mined 0.226min_cos_consec (site) —min_cos_anchor (site) —dataset eurospeechlang daspeaker denmark_20201M042_2020-12-total 65.5schain gain +1.1 dBseam step 1.0 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, balanced body, average recording, quiet background, normally alert
(confusion, contemplation · measured, slightly relaxed, fairly steady, casual) Så det passer jo ikke. (low mumble) Og igen: (low mumble) Jeg synes, det ville være utrolig interessant (ahem) at høre, hvad det er, De Konservative forestiller sig skulle være alternativet.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, contemplation; style: casual, conversational; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 2.6/10; 11.0s, DA.
denmark_20201M042_2020-12-22_0900_29528832_29539840 · in -21.5 dBFS · gain +1.5 dB · eurospeech-00413
(thankfulness gratitude, concentration · normal-paced, slightly relaxed, fairly steady, monologue) den konservative ordfører stiller sig herop og siger, at det er en fuldkommen forkert vej, den her regering har i sin krisepolitik, og man må forstå, at det er helt galt i (low mumble) finansloven. Og så stiller man sig (ahem) efterfølgende op her som (ahem) spørger fra Det Konservative Folkeparti og klandrer regeringen for slet ikke at gøre noget.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, concentration; style: monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.5/10; 15.8s, DA.
denmark_20201M042_2020-12-22_0900_29539840_29555648 · in -20.3 dBFS · gain +0.3 dB · eurospeech-00413
(doubt, confusion, anger · measured, neutral tension, fairly steady, monologue) hvad vil I? Kan vi ikke snart få en retning (ahem) fra (ahem) Det Konservative Folkeparti og finde ud af, (low mumble) hvad jeres svar er? Regeringen har igen investeret (ahem) – som sidste år – i en bedre fremtid (ahem) for børn,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, confusion, anger; style: monologue, casual; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 1.3/10; 16.3s, DA.
denmark_20201M042_2020-12-22_0900_29555648_29571904 · in -21.5 dBFS · gain +1.5 dB · eurospeech-00413
(confusion, intoxication altered states of consciousness · normal-paced, slightly relaxed, moderately variable, casual) bedre pasningsmuligheder for børn, bedre uddannelsesmuligheder for børn og unge. Så det passer ikke, at der ikke sker noget. Kl. 17:12 Anden
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 2.2/10; 12.3s, DA.
denmark_20201M042_2020-12-22_0900_29571904_29584176 · in -23.0 dBFS · gain +3.0 dB · eurospeech-00413
(measured, slightly relaxed, moderately variable, didactic) Nu spurgte jeg ind til de udsatte børn og unge i forhold til (low mumble) statsministerens nytårstale. Men nu er ministeren selv inde på det her omkring minimumsnormeringer.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.3/10; 10.9s, DA.
denmark_20201M042_2020-12-22_0900_29584176_29595120 · in -20.7 dBFS · gain +0.7 dB · eurospeech-00413
Doubt ↓  /  Shameidentity −0.01 emotion 106 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #17

This chain comes from the proxy rule: the same two-sided test as above, but because Shame is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Shame around average — 0.58, higher than 58 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.27.

At the same time Doubt goes the other way, from 0.85 (higher than 85 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.22. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.06 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.91 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.91 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 31 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.902 before conversion and 0.888 after — it fell by 0.014. Neighbour-to-neighbour the worst pair went 0.902 → 0.888. (The earlier render, with segment 1 left raw, scores 0.764 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.274 in the original and +0.292 after conversion — 106 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Doubt, -0.215 became -0.170.

Quality. Mean predicted overall quality across the segments went 3.15 → 3.21 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.902 → 0.888 -0.014identity cos neighbours 0.902 → 0.888d_b rescored +0.274 → +0.292d_a rescored -0.215 → -0.170d_a mined -0.215d_b mined 0.274min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00085_S03610total 30.5schain gain +1.1 dBseam step 1.3 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, slightly dark, fairly smooth, balanced body, no background noise, measured, normally alert, slightly relaxed
(narration, monologue) 坎尼战以后,罗马可谓已陷入绝境,汉尼拔几乎就要实现其征服罗马的梦想了。然而好景不长。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; average recording, no background noise; genuineness 0.5/6; vocal-burst blend 3.8/10; 8.9s, ZH.
ZH_B00085_S03610_W000000 · in -20.6 dBFS · gain +0.6 dB · emolia-02088
(malevolence malice, concentration · formal, monologue) 迦太基是肥尼基人在北非的商业殖民地,大约在公元前九世纪建立,公元前三世纪左右。它是当时地中海西部的强国。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, concentration; style: formal, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 2.5/10; 10.3s, ZH.
ZH_B00085_S03610_W000001 · in -20.4 dBFS · gain +0.4 dB · emolia-02088
(formal, narration) 汉尼巴美占领一个地方,就不得不留一部分兵力守卫。当他要攻打新要塞时,兵力就减少了。他的一部分军队,就是这样零敲碎打的消耗掉了。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 3.8/10; 11.8s, ZH.
ZH_B00085_S03610_W000002 · in -19.9 dBFS · gain -0.1 dB · emolia-02088
Infatuation ↓  /  Emotional Numbnessidentity −0.11 emotion 89 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #18

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness around average — 0.56, higher than 56 % of clips in this corpus — and ends with it strongly present at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.25.

At the same time Infatuation goes the other way, from 0.77 (higher than 77 % of clips in this corpus) to 0.57 (higher than 57 % of clips in this corpus), a change of -0.20. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.97 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.97 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 13 s · en · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.954 before conversion and 0.844 after — it fell by 0.110. Neighbour-to-neighbour the worst pair went 0.954 → 0.844. (The earlier render, with segment 1 left raw, scores 0.486 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.245 in the original and +0.218 after conversion — 89 % of the delta retained, which is most of it. On the other named axis, Infatuation, -0.201 became -0.126.

Quality. Mean predicted overall quality across the segments went 2.83 → 2.99 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.954 → 0.844 -0.110identity cos neighbours 0.954 → 0.844d_b rescored +0.245 → +0.218d_a rescored -0.201 → -0.126d_a mined -0.201d_b mined 0.245min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_PUKyhtbq7cototal 12.1schain gain +1.6 dBseam step 0.4 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(formal, authoritative) Winston Churchill and Zionism Chapelle Manuscript Foundation The Real Churchill' Critical' and a Rebuttal
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 6.9s, EN.
EN_PUKyhtbq7co_W000948 · in -15.8 dBFS · gain -4.2 dB · emolia-01705
(formal, authoritative) A rebuttal to "'The Real Churchill' at the Wayback Machine' Archived 2007-09-12
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.6/10; 5.5s, EN.
EN_PUKyhtbq7co_W000949 · in -15.1 dBFS · gain -5.0 dB · emolia-01705
Infatuation ↓  /  Sournessidentity −0.08 emotion 92 %   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #19

This chain comes from the proxy rule: the same two-sided test as above, but because Sourness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Sourness around average — 0.44, lower than 56 % of clips in this corpus — and ends with it clearly present at 0.66, higher than 66 % of clips in this corpus. That is a total rise of 0.22.

At the same time Infatuation goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.62 (higher than 62 % of clips in this corpus), a change of -0.23. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.22 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 9 s · zh · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.795 before conversion and 0.713 after — it fell by 0.082. Neighbour-to-neighbour the worst pair went 0.795 → 0.713. (The earlier render, with segment 1 left raw, scores 0.708 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sourness moved +0.199 in the original and +0.182 after conversion — 92 % of the delta retained, which is essentially all of it. On the other named axis, Infatuation, -0.234 became -0.333.

Quality. Mean predicted overall quality across the segments went 2.95 → 3.02 (+0.08) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.795 → 0.713 -0.082identity cos neighbours 0.795 → 0.713d_b rescored +0.199 → +0.182d_a rescored -0.234 → -0.333d_a mined -0.234d_b mined 0.217min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00016_S07094total 8.5schain gain +2.4 dBseam step 1.9 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, measured, normally alert
(fairly steady, formal, narration) 即使是一碗饭,一分钱也是不能要的。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 3.7/10; 4.0s, ZH.
ZH_B00016_S07094_W000037 · in -24.7 dBFS · gain +4.7 dB · emolia-03436
(steady, formal, didactic) 既然对利益的追求要服从和符合意义的要求。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, didactic; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.9/10; 4.7s, ZH.
ZH_B00016_S07094_W000038 · in -19.7 dBFS · gain -0.3 dB · emolia-03436
Confusion ↓  /  Contemptidentity +0.17 emotion REVERSED   proxy_taillift__PXR__T0.20__C0.25__INTERNAL · #20

This chain comes from the proxy rule: the same two-sided test as above, but because Contempt is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contempt clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.24.

At the same time Confusion goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.59 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.59 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 9 s · zh · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.538 before conversion and 0.704 after — it rose by 0.165. Neighbour-to-neighbour the worst pair went 0.538 → 0.704. (The earlier render, with segment 1 left raw, scores 0.621 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Contempt moved +0.238 in the original and -0.178 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Confusion, -0.242 became -0.124.

Quality. Mean predicted overall quality across the segments went 2.80 → 2.92 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.538 → 0.704 +0.165identity cos neighbours 0.538 → 0.704d_b rescored +0.238 → -0.178d_a rescored -0.242 → -0.124d_a mined -0.247d_b mined 0.238min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00042_S05331total 8.8schain gain -0.3 dBseam step 0.9 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, no background noise, normally alert, slightly relaxed, fairly steady
(measured, conversational, storytelling) 嗯,你你你你飞哥有多久沒见面?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational, storytelling; average recording, no background noise; genuineness 3.6/6; vocal-burst blend 2.8/10; 3.5s, ZH.
ZH_B00042_S05331_W000035 · in -23.8 dBFS · gain +3.8 dB · emolia-03693
(normal-paced, authoritative, monologue) (low mumble) 好,那我想教一下法师哦,你你这一次说要全面付出,到底你要拼什么?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, monologue; average recording, no background noise; genuineness 2.7/6; vocal-burst blend 2.9/10; 5.5s, ZH.
ZH_B00042_S05331_W000036 · in -21.6 dBFS · gain +1.6 dB · emolia-03693