k-B1-k4 — voice-corrected

B1 at chain length k=4, all corpora, at the mining floor.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_k-B1-k4.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
60segments re-voiced
0.661 → 0.649median worst-to-anchor identity cosine
85 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Hope Enthusiasm Optimism(unconstrained axis: Astonishment Surprise)identity +0.07 emotion 142 %   k-B1-k4 · #1

This chain comes from the one-sided rule: only Hope Enthusiasm Optimism had to get where it was going, by at least 0.50. The other emotion was left completely free.

The chain starts with Hope Enthusiasm Optimism around average — 0.44, lower than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.55.

Nothing was asked of the other axis, and in fact Astonishment Surprise drifts down from 0.99 to 0.36 (-0.63), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.10, then +0.23 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 30 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.659 before conversion and 0.725 after — it rose by 0.066. Neighbour-to-neighbour the worst pair went 0.637 → 0.699. (The earlier render, with segment 1 left raw, scores 0.665 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.549 in the original and +0.778 after conversion — 142 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Astonishment Surprise, -0.631 became -0.554.

Quality. Mean predicted overall quality across the segments went 2.47 → 2.85 (+0.39) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.659 → 0.725 +0.066identity cos neighbours 0.637 → 0.699d_b rescored +0.549 → +0.778d_a rescored -0.631 → -0.554d_a mined -0.631d_b mined 0.549min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_IDE3oxeR6_Etotal 28.9schain gain +4.2 dBseam step 4.1 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · some disfluency
(astonishment surprise, doubt, emotional numbness · normal-paced, energised, fully relaxed, casual) It didn't affect her like it affected me. When I first saw the video, it was like...
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, fully relaxed, moderately variable; timbre is neutral-toned, dark, slightly rough, thin; average clarity, some disfluency, moderate pitch range, audible breath; affect is mildly negative, slightly dominant, neutral openness; reads as astonishment surprise, doubt, emotional numbness; style: casual, storytelling; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 5.5/10; 5.0s, EN.
EN_IDE3oxeR6_E_W000242 · in -21.4 dBFS · gain +1.4 dB · emolia-00784
(astonishment surprise, pleasure ecstasy, triumph · slow, very low-energy, neutral tension, storytelling) The first 20 seconds of the video, I really thought she was alive. I thought she was alive. It blew my mind. I said, oh my god. I said, Zodiacus is alive. I said, oh shit.
full caption & clip details
A middle-aged strongly masculine voice; delivery is very low-energy, slow, neutral tension, volatile; timbre is slightly cool, dark, rough, slightly thin; somewhat unclear, some disfluency, wide pitch range, audible breath; affect is negative, slightly dominant, vulnerable; reads as astonishment surprise, pleasure ecstasy, triumph; style: storytelling, dramatic; below-average recording, quiet background; mildly explicit content; genuineness 3.0/6; vocal-burst blend 5.2/10; 16.2s, EN.
EN_IDE3oxeR6_E_W000243 · in -21.6 dBFS · gain +1.6 dB · emolia-00784
(jealousy and envy, contempt · measured, normally alert, slightly relaxed, casual) Y'all be hatin' on Facebook, sometimes Facebook be doin' some real good stuff.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is warm, dark, slightly rough, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as jealousy and envy, contempt; style: casual, storytelling; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 2.9/10; 4.1s, EN.
EN_IDE3oxeR6_E_W000245 · in -18.3 dBFS · gain -1.7 dB · emolia-00784
(hope enthusiasm optimism, relief, fear · measured, normally alert, slightly relaxed, casual) Yo, for real, I think Facebook is gonna always be here in the future.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as hope enthusiasm optimism, relief, fear; style: casual, monologue; average recording, no background noise; genuineness 2.4/6; vocal-burst blend 2.0/10; 4.2s, EN.
EN_IDE3oxeR6_E_W000246 · in -19.4 dBFS · gain -0.6 dB · emolia-00784
Disgust(unconstrained axis: Interest)identity +0.36 emotion 100 %   k-B1-k4 · #2

This chain comes from the one-sided rule: only Disgust had to get where it was going, by at least 0.50. The other emotion was left completely free.

The chain starts with Disgust below average — 0.33, lower than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.67.

Nothing was asked of the other axis, and in fact Interest drifts down from 1.00 to 0.38 (-0.62), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.23, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.15 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.19 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.15, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 59 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.159 before conversion and 0.524 after — it rose by 0.364. Neighbour-to-neighbour the worst pair went 0.203 → 0.467. (The earlier render, with segment 1 left raw, scores 0.404 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disgust moved +0.671 in the original and +0.670 after conversion — 100 % of the delta retained, which is essentially all of it. On the other named axis, Interest, -0.616 became -0.571.

Quality. Mean predicted overall quality across the segments went 2.80 → 2.96 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.159 → 0.524 +0.364identity cos neighbours 0.203 → 0.467d_b rescored +0.671 → +0.670d_a rescored -0.616 → -0.571d_a mined -0.617d_b mined 0.671min_cos_consec (site) 0.1917min_cos_anchor (site) 0.1543dataset podcastlang enspeaker 681667total 57.8schain gain +2.4 dBseam step 2.4 dBcrossfades 150/150/100 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned
(interest, infatuation, amusement · brisk, energised, neutral tension, casual) his character for Kevin. And you know, the first time we see Kevin, he's like wending his way through the hospital, totally like smoking a cigarette in the hospital. He smokes constantly throughout this movie because you know he's a bad boy writer. But I wasn't feeling anything, and I was just like, what? Where where is where are those old feelings for this this young man I just you know fantasized about for years?
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as interest, infatuation, amusement; style: casual, conversational; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 7.2/10; 24.3s, EN.
681667_00088372 · in -24.1 dBFS · gain +4.1 dB · podcast-06418
(hope enthusiasm optimism · normal-paced, normally alert, neutral tension, casual) And it's not until later in the movie, and we'll get to this, but it's not until later in the movie when the eyes
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism; style: casual, conversational; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 2.9/10; 6.1s, EN.
681667_00090800 · in -24.7 dBFS · gain +4.7 dB · podcast-05972
(pleasure ecstasy, infatuation, hope enthusiasm optimism · slow, normally alert, neutral tension, casual) And it's the it's the I love you more than I ever thought I could love anyone. I want you. I'm smoldering, I'm aflame. It's the eyes, and those come out, and we'll save the we'll save for which character. And it came, they came out, and I was like, there he is. There it is. There's the old stirring.
full caption & clip details
An elderly somewhat feminine voice; delivery is normally alert, slow, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as pleasure ecstasy, infatuation, hope enthusiasm optimism; style: casual, conversational; good recording, quiet background; genuineness 3.0/6; vocal-burst blend 3.6/10; 22.6s, EN.
681667_00091504 · in -24.2 dBFS · gain +4.2 dB · podcast-03030
(disgust, amusement, infatuation · normal-paced, normally alert, slightly relaxed, casual) Love sucks. And so uh (ahem) so he's (ahem) uh he's roommates with Kirby, Emilia's
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as disgust, amusement, infatuation; style: casual, conversational; good recording, no background noise; genuineness 3.6/6; vocal-burst blend 1.1/10; 5.3s, EN.
681667_00095872 · in -27.7 dBFS · gain +7.7 dB · podcast-03944
Hope Enthusiasm Optimism(unconstrained axis: Amusement)identity +0.33 emotion 45 %   k-B1-k4 · #3

This chain comes from the one-sided rule: only Hope Enthusiasm Optimism had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Hope Enthusiasm Optimism around average — 0.55, higher than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.43.

Nothing was asked of the other axis, and in fact Amusement drifts down from 0.98 to 0.78 (-0.21), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.18, then +0.07, then +0.18 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.30 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.34 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.30, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 36 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.270 before conversion and 0.601 after — it rose by 0.331. Neighbour-to-neighbour the worst pair went 0.097 → 0.542. (The earlier render, with segment 1 left raw, scores 0.429 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.428 in the original and +0.193 after conversion — 45 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Amusement, -0.208 became -0.186.

Quality. Mean predicted overall quality across the segments went 2.56 → 2.90 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.270 → 0.601 +0.331identity cos neighbours 0.097 → 0.542d_b rescored +0.428 → +0.193d_a rescored -0.208 → -0.186d_a mined -0.207d_b mined 0.433min_cos_consec (site) 0.3441min_cos_anchor (site) 0.3042dataset podcastlang enspeaker 655741total 34.8schain gain +3.0 dBseam step 0.9 dBcrossfades 100/100/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · some disfluency, wide pitch range, light breath
(amusement, affection, impatience and irritability · brisk, energised, neutral tension, casual) When she looked up at him and said, There's nothing down here for you when he was looking back for J. I just thought she made that show with Jessica. And she never acted before, had she? She was on stage. Oh no, she was in The Secret Life of Bees, that movie, which is one of
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as amusement, affection, impatience and irritability; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 4.9/6; vocal-burst blend 9.3/10; 15.2s, EN.
655741_00206071 · in -20.2 dBFS · gain +0.2 dB · podcast-00296
(thankfulness gratitude, affection, shame · brisk, normally alert, neutral tension, casual) my favorite movies of all time. Jennifer. She should have us on Jennifer. We're common people. Um
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as thankfulness gratitude, affection, shame; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 7.7/10; 4.9s, EN.
655741_00207592 · in -25.7 dBFS · gain +5.7 dB · podcast-04368
(affection, teasing, amusement · normal-paced, normally alert, relaxed, casual) I'm not common. You might be common. We got a lot to say, and we do. We
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is slightly cool, dark, slightly rough, slightly thin; slurred, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as affection, teasing, amusement; style: casual, conversational; below-average recording, some background noise; genuineness 4.4/6; vocal-burst blend 5.1/10; 4.4s, EN.
655741_00208208 · in -21.3 dBFS · gain +1.3 dB · podcast-04360
(hope enthusiasm optimism, contentment, elation · brisk, energised, neutral tension, casual) call. Yes, uh, reach out, reach out through zagarrealty.com. We can be found there. And (low mumble) uh and that too, we always forget. If you're looking to buy or sell a home (ahem) uh in five states around here, or build give
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism, contentment, elation; style: casual, conversational; below-average recording, some background noise; genuineness 4.0/6; vocal-burst blend 9.6/10; 10.8s, EN.
655741_00210136 · in -22.2 dBFS · gain +2.2 dB · podcast-04360
Hope Enthusiasm Optimism(unconstrained axis: Doubt)identity +0.47 emotion 94 %   k-B1-k4 · #4

This chain comes from the one-sided rule: only Hope Enthusiasm Optimism had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Hope Enthusiasm Optimism clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.37.

Nothing was asked of the other axis, and in fact Doubt drifts down from 0.93 to 0.05 (-0.87), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.18, then +0.10, then +0.09 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 36 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.189 before conversion and 0.658 after — it rose by 0.470. Neighbour-to-neighbour the worst pair went 0.185 → 0.658. (The earlier render, with segment 1 left raw, scores 0.480 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.366 in the original and +0.344 after conversion — 94 % of the delta retained, which is essentially all of it. On the other named axis, Doubt, -0.872 became -0.275.

Quality. Mean predicted overall quality across the segments went 2.71 → 2.90 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.189 → 0.658 +0.470identity cos neighbours 0.185 → 0.658d_b rescored +0.366 → +0.344d_a rescored -0.872 → -0.275d_a mined -0.871d_b mined 0.366min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_ExjyhzglGtAtotal 35.2schain gain +5.0 dBseam step 0.7 dBcrossfades 150/150/150 ms
Script — 4 chunks, 4 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(doubt, contemplation · slightly relaxed, some disfluency, average clarity, conversational) Are there any questions that (low mumble) have come about that you feel I should answer right now?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as doubt, contemplation; style: conversational, monologue; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.3/10; 5.2s, EN.
EN_ExjyhzglGtA_W000515 · in -18.9 dBFS · gain -1.1 dB · emolia-01308
(slightly relaxed, some disfluency, average clarity, casual) So Jeff, right now we are good to go. You can go ahead and (low mumble) proceed forward with the (low mumble) quiz.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 1.9/10; 5.8s, EN.
EN_ExjyhzglGtA_W000516 · in -14.3 dBFS · gain -5.7 dB · emolia-01308
(slightly relaxed, frequent disfluency, average clarity, casual) Again, if you have them and you think of some, (ahem) uh, as we play this quiz, please put them in there and we will, (low mumble) uhm, we will have all of them.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 1.8/10; 7.7s, EN.
EN_ExjyhzglGtA_W000517 · in -19.5 dBFS · gain -0.5 dB · emolia-01308
(hope enthusiasm optimism, fatigue exhaustion, triumph · neutral tension, some disfluency, somewhat unclear, casual) Uh, (low mumble) finally, in Google Classroom and in the follow-ups, I'm going to be sending out both, after Thursday, once we do the math, uh, (low mumble) I will send out a follow-up survey. But even right now, (low mumble) uhm, after this, I'm going to be sending an assignment in Google Classroom. You don't have to do it.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism, fatigue exhaustion, triumph; style: casual, monologue; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 6.3/10; 17.1s, EN.
EN_ExjyhzglGtA_W000518 · in -18.6 dBFS · gain -1.4 dB · emolia-01308
Interest(unconstrained axis: Emotional Numbness)identity −0.03 emotion 113 %   k-B1-k4 · #5

This chain comes from the one-sided rule: only Interest had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Interest around average — 0.48, lower than 52 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.48.

Nothing was asked of the other axis, and in fact Emotional Numbness drifts down from 0.82 to 0.32 (-0.49), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.18, then +0.18, then +0.12 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.85 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.85 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 47 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.693 before conversion and 0.663 after — it fell by 0.029. Neighbour-to-neighbour the worst pair went 0.802 → 0.775. (The earlier render, with segment 1 left raw, scores 0.605 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.481 in the original and +0.541 after conversion — 113 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, -0.495 became -0.385.

Quality. Mean predicted overall quality across the segments went 2.90 → 3.10 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.693 → 0.663 -0.029identity cos neighbours 0.802 → 0.775d_b rescored +0.481 → +0.541d_a rescored -0.495 → -0.385d_a mined -0.495d_b mined 0.480min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_D2GT9zGhAPUtotal 45.8schain gain +3.6 dBseam step 2.3 dBcrossfades 150/150/100 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, balanced body, slightly relaxed, average clarity, moderate pitch range
(normal-paced, normally alert, moderately variable, casual) In flight, the wings carry all the weight. So if we put the, we distribute that weight out on the wing,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: casual, conversational; good recording, quiet background; genuineness 3.4/6; vocal-burst blend 4.9/10; 5.9s, EN.
EN_D2GT9zGhAPU_W000035 · in -16.5 dBFS · gain -3.5 dB · emolia-02536
(emotional numbness, concentration · measured, normally alert, fairly steady, monologue) We don't need to carry it through other things. So we have a lighter fuselage structure with, (low mumble) uhm, with engines under the wing than we do otherwise.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness, concentration; style: monologue, casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.7/10; 8.5s, EN.
EN_D2GT9zGhAPU_W000036 · in -22.2 dBFS · gain +2.2 dB · emolia-02536
(fear, teasing · normal-paced, normally alert, fairly steady, casual) (ahem) Uhm, easier terminal servicing? Nope. We've got to lift the aircraft up higher off the ground. That means you need steps and ladders where you might not need it otherwise. (ahem) Uhm, for things like getting passengers on and off, fueling the aircraft, (ahem) uhm, dumping the lavs, and remember the blue juice is not
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as fear, teasing; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 3.9/10; 16.4s, EN.
EN_D2GT9zGhAPU_W000037 · in -19.3 dBFS · gain -0.7 dB · emolia-02536
(interest, concentration, elation · brisk, energised, fairly steady, conversational) And worse weight and balance. No, actually it's really nice for weight and balance. You don't have weird concentrations of weight at the extremes of your fuselage. Engines are really heavy and dense. So we stick an engine way at the back of the fuselage, like on the 727.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as interest, concentration, elation; style: conversational, casual; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 4.6/10; 15.6s, EN.
EN_D2GT9zGhAPU_W000039 · in -18.6 dBFS · gain -1.4 dB · emolia-02536
Contentment(unconstrained axis: Shame)identity −0.00 emotion 16 %   k-B1-k4 · #6

This chain comes from the one-sided rule: only Contentment had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Contentment clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Shame drifts down from 0.99 to 0.84 (-0.15), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.16, then -0.08, then +0.16 — not a clean run: step 2 moves back the other way by 0.08 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 67 s · portuguese · mls

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.912 before conversion and 0.910 after — it fell by 0.001. Neighbour-to-neighbour the worst pair went 0.914 → 0.926. (The earlier render, with segment 1 left raw, scores 0.751 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.241 in the original and +0.039 after conversion — 16 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Shame, -0.148 became -0.096.

Quality. Mean predicted overall quality across the segments went 3.23 → 3.44 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.912 → 0.910 -0.001identity cos neighbours 0.914 → 0.926d_b rescored +0.241 → +0.039d_a rescored -0.148 → -0.096d_a mined -0.148d_b mined 0.242min_cos_consec (site) 0.9281min_cos_anchor (site) 0.9225dataset mlslang portuguesespeaker 9351total 65.6schain gain +3.0 dBseam step 1.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a middle-aged somewhat masculine voice · balanced body, average recording, quiet background, slightly relaxed
(shame, disappointment, anger · measured, subdued, fairly steady, whispered) como sabe há muitos desgostos contra o regente se o imperador já tivesse a idade de constituição é que era bom ia-se embora o regente e o resto pois é verdade creio que sim entretanto nunca tinha pensado nisto seriamente mas as cousas são assim mesmo que acha
full caption & clip details
A middle-aged somewhat masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is slightly cool, dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as shame, disappointment, anger; style: whispered, monologue; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 3.3/10; 17.8s, PORTUGUESE.
9351_9163_000345 · in -27.2 dBFS · gain +7.2 dB · mls-00124
(disappointment · measured, subdued, fairly steady, monologue) acho que fez bem em todo o caso peco lhe segredo não diga nada a mamãe crê que ela se oponha não mas pode ser que não se alcance nada e para lhe não dar uma esperança que pode falhar e só isto era plausível a explicação
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment; style: monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 3.9/10; 17.2s, PORTUGUESE.
9351_9163_000242 · in -28.4 dBFS · gain +8.4 dB · mls-00123
(normal-paced, normally alert, fairly steady, monologue) explicação prometi lhe não dizer nada creio que falamos ainda de política e da política daqueles últimos dez anos que não era pouca nem plácida félix não tinha certamente um plano de idéias e apreciações originais
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 3.0/10; 13.4s, PORTUGUESE.
9351_9163_000142 · in -29.2 dBFS · gain +9.2 dB · mls-00123
(contentment, thankfulness gratitude · normal-paced, normally alert, steady, monologue) através das palavras dele apalpava eu as fórmulas e os juízos do círculo ou das pessoas com quem ele lidava para o fim de encetar a carreira agora a particularidade dele era a clareza e retidão de espírito precisas para só recolher do que ouvia a parte sã e justa
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as contentment, thankfulness gratitude; style: monologue, narration; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 3.8/10; 17.8s, PORTUGUESE.
9351_9163_000023 · in -27.5 dBFS · gain +7.5 dB · mls-00123
Pain(unconstrained axis: Distress)identity +0.07 emotion 44 %   k-B1-k4 · #7

This chain comes from the one-sided rule: only Pain had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Pain below average — 0.35, lower than 65 % of clips in this corpus — and ends with it strongly present at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.45.

Nothing was asked of the other axis, and in fact Distress drifts down from 0.92 to 0.44 (-0.48), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.17, then +0.08 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.90 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.90 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 31 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.825 before conversion and 0.899 after — it rose by 0.074. Neighbour-to-neighbour the worst pair went 0.833 → 0.890. (The earlier render, with segment 1 left raw, scores 0.807 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.449 in the original and +0.196 after conversion — 44 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Distress, -0.477 became +0.000.

Quality. Mean predicted overall quality across the segments went 2.97 → 3.24 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.825 → 0.899 +0.074identity cos neighbours 0.833 → 0.890d_b rescored +0.449 → +0.196d_a rescored -0.477 → +0.000d_a mined -0.478d_b mined 0.449min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00052_S03565total 29.9schain gain +0.7 dBseam step 0.7 dBcrossfades 150/100/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, no background noise, normal-paced, normally alert
(distress · fairly steady, monologue, formal) 如果按照目前的词汇来说,亚特兰蒂斯最初只是利莫里亚文明的一块殖民地或新开发的一片大陆。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as distress; style: monologue, formal; average recording, no background noise; genuineness 0.7/6; vocal-burst blend 2.5/10; 8.4s, ZH.
ZH_B00052_S03565_W000001 · in -12.9 dBFS · gain -7.0 dB · emolia-03797
(triumph, hope enthusiasm optimism · fairly steady, formal, narration) 后来,由于该地区的主流信仰发生了变化,从均衡的唯心与唯物思想转变为更渴求物质刺激,逐渐偏向了唯物科学。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, hope enthusiasm optimism; style: formal, narration; average recording, no background noise; genuineness 0.4/6; vocal-burst blend 3.0/10; 9.7s, ZH.
ZH_B00052_S03565_W000002 · in -13.2 dBFS · gain -6.8 dB · emolia-03797
(fairly steady, formal, monologue) 但从地球整个星球层面来说,这两个文明毕竟属于同根同源。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 2.3/10; 5.0s, ZH.
ZH_B00052_S03565_W000003 · in -13.8 dBFS · gain -6.2 dB · emolia-03797
(steady, monologue, formal) 在面对旧帝国与大洪水这种共同的敌人与灾难面前,双方自然能够冰释前嫌与同舟共济。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 0.8/6; vocal-burst blend 2.1/10; 7.3s, ZH.
ZH_B00052_S03565_W000004 · in -13.6 dBFS · gain -6.4 dB · emolia-03797
Astonishment Surprise(unconstrained axis: Disappointment)identity −0.07 emotion 103 %   k-B1-k4 · #8

This chain comes from the one-sided rule: only Astonishment Surprise had to get where it was going, by at least 0.50. The other emotion was left completely free.

The chain starts with Astonishment Surprise below average — 0.36, lower than 64 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.50.

Nothing was asked of the other axis, and in fact Disappointment drifts down from 0.90 to 0.39 (-0.50), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.16, then +0.13, then +0.20 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 34 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.913 before conversion and 0.848 after — it fell by 0.065. Neighbour-to-neighbour the worst pair went 0.920 → 0.844. (The earlier render, with segment 1 left raw, scores 0.716 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.505 in the original and +0.520 after conversion — 103 % of the delta retained, which is essentially all of it. On the other named axis, Disappointment, -0.501 became -0.517.

Quality. Mean predicted overall quality across the segments went 2.95 → 3.11 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.913 → 0.848 -0.065identity cos neighbours 0.920 → 0.844d_b rescored +0.505 → +0.520d_a rescored -0.501 → -0.517d_a mined -0.501d_b mined 0.504min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00074_S07774total 32.9schain gain +2.3 dBseam step 0.8 dBcrossfades 150/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(steady, formal, newsreading) Plymouth's colonists faced great hardships and earned few profits for their investors, who sold their interests to them in 1627
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.3/10; 7.7s, EN.
EN_B00074_S07774_W000008 · in -17.0 dBFS · gain -3.0 dB · emolia-01661
(sadness, helplessness, disappointment · steady, formal, narration) Bridges were fairly uncommon, since they were expensive to maintain, and fines were imposed
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sadness, helplessness, disappointment; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.5/10; 7.9s, EN.
EN_B00074_S07774_W000009 · in -15.9 dBFS · gain -4.1 dB · emolia-01661
(fairly steady, formal, newsreading) The proceedings were arranged so that the time had expired for the colonial authorities to defend the charter, before they even learned of the event
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.9/10; 7.3s, EN.
EN_B00074_S07774_W000010 · in -17.1 dBFS · gain -3.0 dB · emolia-01661
(fairly steady, formal, newsreading) By 1660, the colony's merchant fleet was estimated at 200 ships and, by the end of the century, its shipyards were estimated to turn out several hundred ships annually
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.4/10; 10.6s, EN.
EN_B00074_S07774_W000011 · in -18.1 dBFS · gain -1.9 dB · emolia-01661
Jealousy and Envy(unconstrained axis: Embarrassment)identity +0.44 emotion REVERSED   k-B1-k4 · #9

This chain comes from the one-sided rule: only Jealousy and Envy had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Jealousy and Envy clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.30.

Nothing was asked of the other axis, and in fact Embarrassment barely moves at all, sitting near 1.00 throughout.

It takes 4 clips to get there. Clip to clip the moves are +0.21, then -0.15, then +0.24 — not a clean run: step 2 moves back the other way by 0.15 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.14 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.23 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.14, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 39 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.045 before conversion and 0.480 after — it rose by 0.436. Neighbour-to-neighbour the worst pair went 0.106 → 0.496. (The earlier render, with segment 1 left raw, scores 0.200 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

The emotional move did not survive. Re-scored end to end, Jealousy and Envy moved +0.304 in the original and -0.517 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Embarrassment, -0.049 became -0.117.

Quality. Mean predicted overall quality across the segments went 2.68 → 2.89 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.045 → 0.480 +0.436identity cos neighbours 0.106 → 0.496d_b rescored +0.304 → -0.517d_a rescored -0.049 → -0.117d_a mined -0.049d_b mined 0.305min_cos_consec (site) 0.2279min_cos_anchor (site) 0.1352dataset podcastlang enspeaker 224158total 37.8schain gain +3.9 dBseam step 3.2 dBcrossfades 150/100/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, normally alert, light breath
(embarrassment, shame, amusement · normal-paced, neutral tension, moderately variable, casual) I mean, yeah. (ahem) At least he didn't try to hide it and like cover for him. I mean yeah, I guess that is better. Yeah. Seth Rogan's actually making a (ahem) uh CGI Ninja Turtle movie currently.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as embarrassment, shame, amusement; style: casual, conversational; average recording, some background noise; mildly explicit content; genuineness 5.5/6; vocal-burst blend 6.3/10; 16.1s, EN.
224158_00031528 · in -26.7 dBFS · gain +6.7 dB · podcast-01181
(sexual lust, teasing, doubt · measured, slightly relaxed, fairly steady, casual) Wait, did you say he's making a CGI Ninja Turtle movie?
full caption & clip details
A child masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, balanced body; slurred, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as sexual lust, teasing, doubt; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 5.8/10; 3.6s, EN.
224158_00033264 · in -28.5 dBFS · gain +8.5 dB · podcast-01187
(astonishment surprise, intoxication altered states of consciousness, confusion · normal-paced, neutral tension, moderately variable, casual) Dude, the newer, not the newer ones, but like the most recent live action Ninja Total movies were not really that good. (wistful sigh) They could have gotten a better, they could have gotten a better person to play April instead of (ahem) um not April, God what is her name? Was it April?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as astonishment surprise, intoxication altered states of consciousness, confusion; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.5/6; vocal-burst blend 6.8/10; 15.3s, EN.
224158_00033840 · in -24.7 dBFS · gain +4.7 dB · podcast-01185
(jealousy and envy, embarrassment, pride · normal-paced, slightly relaxed, fairly steady, casual) it's oh yeah, they could have gotten like somebody better than Megan Fox to play her.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as jealousy and envy, embarrassment, pride; style: casual, conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 6.7/10; 3.2s, EN.
224158_00035496 · in -24.8 dBFS · gain +4.8 dB · podcast-01195
Infatuation(unconstrained axis: Amusement)identity +0.06 emotion 55 %   k-B1-k4 · #10

This chain comes from the one-sided rule: only Infatuation had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Infatuation around average — 0.51, right about the corpus median — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.45.

Nothing was asked of the other axis, and in fact Amusement drifts down from 1.00 to 0.83 (-0.17), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.21, then +0.21, then +0.03 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.24 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.24 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.24, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 35 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.374 before conversion and 0.433 after — it rose by 0.059. Neighbour-to-neighbour the worst pair went 0.374 → 0.433. (The earlier render, with segment 1 left raw, scores 0.360 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.447 in the original and +0.248 after conversion — 55 % of the delta retained. On the other named axis, Amusement, -0.167 became -0.152.

Quality. Mean predicted overall quality across the segments went 2.62 → 3.04 (+0.43) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.374 → 0.433 +0.059identity cos neighbours 0.374 → 0.433d_b rescored +0.447 → +0.248d_a rescored -0.167 → -0.152d_a mined -0.166d_b mined 0.447min_cos_consec (site) 0.2406min_cos_anchor (site) 0.2406dataset podcastlang enspeaker 833834total 34.2schain gain +3.6 dBseam step 0.9 dBcrossfades 150/150/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, normal-paced
(amusement, intoxication altered states of consciousness, pleasure ecstasy · normally alert, fully relaxed, volatile, casual) (chuckle) Uh words of affirmation.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, volatile; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, frequent disfluency, wide pitch range, breathless; affect is positive, slightly submissive, neutral openness; reads as amusement, intoxication altered states of consciousness, pleasure ecstasy; style: casual, playful; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 0.0/10; 3.1s, EN.
833834_00226040 · in -22.5 dBFS · gain +2.5 dB · podcast-00365
(amusement, teasing, malevolence malice · normally alert, neutral tension, moderately variable, casual) like they just they're like what's the analogy if you put your wife and your dog in the boot of the car and drive around the block for half an hour and then open the boot, which one's happy to send
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as amusement, teasing, malevolence malice; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 6.0/6; vocal-burst blend 7.0/10; 7.9s, EN.
833834_00231567 · in -22.8 dBFS · gain +2.8 dB · podcast-01526
(contentment, interest, infatuation · very low-energy, relaxed, fairly steady, casual) know. Four. That's (ahem)
full caption & clip details
An adult masculine voice; delivery is very low-energy, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as contentment, interest, infatuation; style: casual, conversational; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 3.3/10; 17.2s, EN.
833834_00234871 · in -26.6 dBFS · gain +6.6 dB · podcast-01622
(infatuation, embarrassment, pleasure ecstasy · normally alert, relaxed, fairly steady, casual) go for that. I was gonna say mine, because it's so funny now that I hear them again. (low mumble) Um mine's acts of
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as infatuation, embarrassment, pleasure ecstasy; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 2.7/10; 6.6s, EN.
833834_00239184 · in -23.7 dBFS · gain +3.7 dB · podcast-01530
Doubt(unconstrained axis: Concentration)identity +0.46 emotion 92 %   k-B1-k4 · #11

This chain comes from the one-sided rule: only Doubt had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Doubt clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.33.

Nothing was asked of the other axis, and in fact Concentration drifts down from 0.96 to 0.81 (-0.15), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.24, then +0.09 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 43 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.119 before conversion and 0.574 after — it rose by 0.455. Neighbour-to-neighbour the worst pair went 0.196 → 0.640. (The earlier render, with segment 1 left raw, scores 0.519 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.379 in the original and +0.348 after conversion — 92 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.151 became -0.084.

Quality. Mean predicted overall quality across the segments went 2.94 → 3.15 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.119 → 0.574 +0.455identity cos neighbours 0.196 → 0.640d_b rescored +0.379 → +0.348d_a rescored -0.151 → -0.084d_a mined -0.151d_b mined 0.334min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_4P6nD-vlLx0total 42.5schain gain -0.2 dBseam step 2.4 dBcrossfades 100/150/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-toned, balanced body, slightly relaxed, fairly steady
(concentration · slow, very low-energy, frequent disfluency, monologue) Uh, (low mumble) share data and communicate over distances, internationally, uh, (low mumble) research networks, et cetera. And more and more we see that the usage is also happening locally.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 0.0/10; 11.8s, EN.
EN_4P6nD-vlLx0_W000082 · in -20.2 dBFS · gain +0.2 dB · emolia-02574
(concentration · slow, very low-energy, frequent disfluency, monologue) And this is where this next group of users can really benefit from the internet. That was first to connect from large distances. Also locally you can help organize yourself, (low mumble) ranging from (low mumble) doing business to, (low mumble) uh,
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 0.0/10; 16.3s, EN.
EN_4P6nD-vlLx0_W000083 · in -24.2 dBFS · gain +4.2 dB · emolia-02574
(contemplation · measured, normally alert, some disfluency, monologue) Facilitating this internet to also serve people who don't use Latin script, who don't use the English language is (ahem) for me such an important thing.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation; style: monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.5/10; 8.5s, EN.
EN_4P6nD-vlLx0_W000085 · in -21.6 dBFS · gain +1.6 dB · emolia-02574
(doubt · normal-paced, normally alert, little disfluency, monologue) So if we want the internet then to be available in different languages, what is needed to achieve this?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as doubt; style: monologue, casual; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.0/10; 6.5s, EN.
EN_4P6nD-vlLx0_W000086 · in -20.0 dBFS · gain +0.0 dB · emolia-02574
Disgust(unconstrained axis: Longing)identity −0.02 emotion 189 %   k-B1-k4 · #12

This chain comes from the one-sided rule: only Disgust had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Disgust clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.34.

Nothing was asked of the other axis, and in fact Longing drifts down from 1.00 to 0.77 (-0.23), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.16, then +0.04, then +0.14 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.74 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.74 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 42 s · de · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.662 before conversion and 0.640 after — it fell by 0.023. Neighbour-to-neighbour the worst pair went 0.714 → 0.711. (The earlier render, with segment 1 left raw, scores 0.664 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disgust moved +0.337 in the original and +0.636 after conversion — 189 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Longing, -0.226 became -0.168.

Quality. Mean predicted overall quality across the segments went 2.93 → 2.98 (+0.05) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.662 → 0.640 -0.023identity cos neighbours 0.714 → 0.711d_b rescored +0.337 → +0.636d_a rescored -0.226 → -0.168d_a mined -0.226d_b mined 0.337min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_TRC5RUEl2qytotal 40.4schain gain +1.1 dBseam step 1.8 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, slightly relaxed, clear, light breath
(longing, teasing, jealousy and envy · normal-paced, normally alert, fairly steady, storytelling) Er verlor ein bisschen den Mut, aber strengte sich doch noch einmal an.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as longing, teasing, jealousy and envy; style: storytelling, formal; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.5/10; 3.1s, DE.
DE_TRC5RUEl2qy_W000032 · in -17.4 dBFS · gain -2.6 dB · emolia-00252
(awe, longing, infatuation · normal-paced, normally alert, fairly steady, ASMR) Es wird fröhlich sein. Auch ich werde die Sterne betrachten. Alle Sterne werden wie Brunnen sein mit einer verrosteten Winde. Alle Sterne werden mir zu trinken geben. Ich schwieg.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as awe, longing, infatuation; style: ASMR, narration; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.4/10; 10.8s, DE.
DE_TRC5RUEl2qy_W000033 · in -18.4 dBFS · gain -1.6 dB · emolia-00252
(disappointment, anger, jealousy and envy · measured, very low-energy, fairly steady, narration) Es wird so lustig sein. Du wirst 500 Millionen Glöckchen haben. Ich werde 500 Millionen Brunnen haben. Dann schwieg auch er, denn er weinte. Dort ist es. Lass mich allein einen Schritt machen. Und er setzte sich hin, denn er hatte Angst. Er sagte, weißt du, meine Blume, ich bin für sie verantwortlich und sie ist so schwach.
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, fairly narrow pitch, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as disappointment, anger, jealousy and envy; style: narration, ASMR; good recording, quiet background; genuineness 1.6/6; vocal-burst blend 2.6/10; 22.3s, DE.
DE_TRC5RUEl2qy_W000034 · in -18.9 dBFS · gain -1.1 dB · emolia-00252
(disgust, jealousy and envy, sourness · normal-paced, normally alert, steady, formal) Sie ist so leichtgläubig. Sie hat nur vier lächerliche Dornen, um sich gegen die Welt zu schützen.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as disgust, jealousy and envy, sourness; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 4.9s, DE.
DE_TRC5RUEl2qy_W000035 · in -14.2 dBFS · gain -5.8 dB · emolia-00252
Concentration(unconstrained axis: Malevolence Malice)identity +0.01 emotion 81 %   k-B1-k4 · #13

This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Concentration clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.36.

Nothing was asked of the other axis, and in fact Malevolence Malice drifts down from 0.78 to 0.42 (-0.36), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.21, then +0.02, then +0.13 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 52 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.794 before conversion and 0.805 after — it rose by 0.011. Neighbour-to-neighbour the worst pair went 0.885 → 0.872. (The earlier render, with segment 1 left raw, scores 0.797 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.363 in the original and +0.294 after conversion — 81 % of the delta retained, which is most of it. On the other named axis, Malevolence Malice, -0.365 became -0.440.

Quality. Mean predicted overall quality across the segments went 3.09 → 3.24 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.794 → 0.805 +0.011identity cos neighbours 0.885 → 0.872d_b rescored +0.363 → +0.294d_a rescored -0.365 → -0.440d_a mined -0.365d_b mined 0.362min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00008_S04771total 50.8schain gain +1.4 dBseam step 0.5 dBcrossfades 150/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, normally alert, slightly relaxed, fairly steady, some disfluency, clear
(measured, monologue, whispered) 第二,产学研一体化,鼓励科技创新和转化应用。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: monologue, whispered; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 3.5/10; 5.3s, ZH.
ZH_B00008_S04771_W000034 · in -20.3 dBFS · gain +0.3 dB · emolia-03353
(measured, didactic, whispered) 第三,松绑房地产调控,促进软着陆。第四,繁荣资本市场,提振市场信心,提高居民财富效应,促进科技创新。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, whispered; average recording, no background noise; genuineness 0.8/6; vocal-burst blend 2.0/10; 11.1s, ZH.
ZH_B00008_S04771_W000035 · in -21.9 dBFS · gain +1.9 dB · emolia-03353
(measured, didactic, monologue) 第五,进一步扩大REITS试点范围,盘活存量资产。第六,大力实施都市圈城市群战略。第七,完善生育政策体系,给予育儿优惠措施。第八,促进财务转型。从土地财政转型,股权财政。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 3.3/10; 19.3s, ZH.
ZH_B00008_S04771_W000036 · in -21.3 dBFS · gain +1.3 dB · emolia-03353
(concentration · normal-paced, didactic, monologue) 如果措施有力全力拼经济,二零二三年中国经济有望重新引领全球,预计全球三大经济体欧洲经济延续衰退。美国经济从滞胀步入衰退,中国经济从筑底步入复苏。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: didactic, monologue; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 2.7/10; 15.6s, ZH.
ZH_B00008_S04771_W000037 · in -20.9 dBFS · gain +0.9 dB · emolia-03353
Emotional Numbness(unconstrained axis: Infatuation)identity −0.06 emotion 55 %   k-B1-k4 · #14

This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Emotional Numbness clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.28.

Nothing was asked of the other axis, and in fact Infatuation drifts down from 0.80 to 0.01 (-0.79), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.22, then -0.02, then +0.08 — not a clean run: step 2 moves back the other way by 0.02 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.92 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.92 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 43 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.916 before conversion and 0.861 after — it fell by 0.056. Neighbour-to-neighbour the worst pair went 0.936 → 0.893. (The earlier render, with segment 1 left raw, scores 0.777 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.282 in the original and +0.154 after conversion — 55 % of the delta retained. On the other named axis, Infatuation, -0.792 became -0.378.

Quality. Mean predicted overall quality across the segments went 2.95 → 3.09 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.916 → 0.861 -0.056identity cos neighbours 0.936 → 0.893d_b rescored +0.282 → +0.154d_a rescored -0.792 → -0.378d_a mined -0.792d_b mined 0.283min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_adhbxa7VMgctotal 41.5schain gain +1.3 dBseam step 0.6 dBcrossfades 150/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-bright, fairly smooth, good recording, measured, normally alert, slightly relaxed, steady, no disfluency
(fairly narrow pitch, light breath, formal, monologue) Indeed, hyperbolic space can have more than two dimensions and one can
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.3/10; 6.8s, EN.
EN_adhbxa7VMgc_W000036 · in -16.0 dBFS · gain -4.0 dB · emolia-00373
(emotional numbness · moderate pitch range, light breath, formal, casual) Copies of hyperbolic space to get higher dimensional models of anti-de-sitter space.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, casual; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.4/10; 6.3s, EN.
EN_adhbxa7VMgc_W000037 · in -15.0 dBFS · gain -5.0 dB · emolia-00373
(emotional numbness · fairly narrow pitch, light breath, formal, monologue) An important feature of anti-de-sitter space is its boundary, which looks like a cylinder in the case of three-dimensional anti-de-sitter space.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 9.9s, EN.
EN_adhbxa7VMgc_W000038 · in -15.1 dBFS · gain -4.9 dB · emolia-00373
(emotional numbness, awe · fairly narrow pitch, no audible breath, formal, newsreading) is given by the boundary of anti-de Sitter space. This observation is the starting point for AdS.CFT correspondence, which states that the boundary of anti-de Sitter space can be regarded as the
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, no disfluency, fairly narrow pitch, no audible breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, awe; style: formal, newsreading; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.0/10; 19.1s, EN.
EN_adhbxa7VMgc_W000039 · in -15.9 dBFS · gain -4.1 dB · emolia-00373
Disgust(unconstrained axis: Intoxication Altered States of Consciousness)identity +0.01 emotion REVERSED   k-B1-k4 · #15

This chain comes from the one-sided rule: only Disgust had to get where it was going, by at least 0.50. The other emotion was left completely free.

The chain starts with Disgust below average — 0.33, lower than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.61.

Nothing was asked of the other axis, and in fact Intoxication Altered States of Consciousness drifts down from 0.95 to 0.89 (-0.06), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.23, then +0.17 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 56 s · da · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.921 before conversion and 0.935 after — it rose by 0.014. Neighbour-to-neighbour the worst pair went 0.921 → 0.927. (The earlier render, with segment 1 left raw, scores 0.783 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The emotional move did not survive. Re-scored end to end, Disgust moved +0.607 in the original and -0.120 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Intoxication Altered States of Consciousness, -0.060 became -0.043.

Quality. Mean predicted overall quality across the segments went 3.08 → 3.36 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.921 → 0.935 +0.014identity cos neighbours 0.921 → 0.927d_b rescored +0.607 → -0.120d_a rescored -0.060 → -0.043d_a mined -0.060d_b mined 0.607min_cos_consec (site) —min_cos_anchor (site) —dataset eurospeechlang daspeaker denmark_20141M086_2015-05-total 55.4schain gain +0.7 dBseam step 0.8 dBcrossfades 150/150/150 ms
Script — 4 chunks, 4 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-toned, slightly dark, slightly rough, balanced body, average recording, somewhat unclear, audible breath
(intoxication altered states of consciousness, doubt · measured, very low-energy, neutral tension, monologue) For det første skal vi have de 200 af (low mumble) AKT-vognene væk fra alle bevogtningsopgaverne, og (ahem) jeg tror, at hvis vi laver en statistik på, hvor tit man har været ude at se efter stjålne ting, er det lig nul.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness, doubt; style: monologue, casual; average recording, some background noise; genuineness 4.6/6; vocal-burst blend 3.5/10; 16.1s, DA.
denmark_20141M086_2015-05-05_1300_15113792_15129856 · in -23.1 dBFS · gain +3.1 dB · eurospeech-00243
(sourness, fatigue exhaustion · measured, very low-energy, neutral tension, monologue) Men (low mumble) så er spørgsmålet også, hvis man gjorde det, hvor meget det havde hjulpet, for jeg tror, at i hvert fald de, der er udlændinge, forsvinder (low mumble) med tingene (low mumble) meget,
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as sourness, fatigue exhaustion; style: monologue, casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 3.7/10; 12.6s, DA.
denmark_20141M086_2015-05-05_1300_15129856_15142432 · in -22.2 dBFS · gain +2.2 dB · eurospeech-00243
(concentration · measured, normally alert, slightly relaxed, monologue) (low mumble) eller det bliver omsmeltet, (low mumble) når man går efter guld- og sølvsmykker, og så er der også endelig det ved det, at desværre (ahem) er der ikke helt styr på den slags ting blandt dem, der får dem stjålet.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue; average recording, some background noise; genuineness 4.0/6; vocal-burst blend 2.4/10; 14.2s, DA.
denmark_20141M086_2015-05-05_1300_15142432_15156672 · in -23.8 dBFS · gain +3.8 dB · eurospeech-00243
(disgust, bitterness, concentration · normal-paced, normally alert, slightly relaxed, monologue) Så jeg vil godt opfordre folk, der vil have deres (low mumble) ting igen – det er også en måde, man kan opklare sager på – til at være bedre til at registrere, hvad man har, og være bedre til at få det efterlyst, således at politiet, når de ude på
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as disgust, bitterness, concentration; style: monologue, authoritative; average recording, some background noise; genuineness 2.3/6; vocal-burst blend 2.6/10; 13.1s, DA.
denmark_20141M086_2015-05-05_1300_15156672_15169776 · in -24.2 dBFS · gain +4.2 dB · eurospeech-00243
Longing(unconstrained axis: Fear)identity −0.04 emotion 132 %   k-B1-k4 · #16

This chain comes from the one-sided rule: only Longing had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Longing clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.31.

Nothing was asked of the other axis, and in fact Fear drifts down from 1.00 to 0.45 (-0.55), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.19, then +0.11, then +0.00 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.83 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.83 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 46 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.538 before conversion and 0.499 after — it fell by 0.039. Neighbour-to-neighbour the worst pair went 0.538 → 0.499. (The earlier render, with segment 1 left raw, scores 0.389 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.306 in the original and +0.402 after conversion — 132 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Fear, -0.553 became -0.865.

Quality. Mean predicted overall quality across the segments went 2.81 → 2.90 (+0.09) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.538 → 0.499 -0.039identity cos neighbours 0.538 → 0.499d_b rescored +0.306 → +0.402d_a rescored -0.553 → -0.865d_a mined -0.553d_b mined 0.306min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00051_S02556total 44.8schain gain -1.8 dBseam step 2.3 dBcrossfades 150/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, average recording, slightly relaxed, fairly steady, clear, light breath
(fear, disappointment, awe · normal-paced, subdued, almost no disfluency, whispered) But when he noticed with gentle concern that Peter did not seem to know that this was rather an odd way of getting your bread and butter, nor even that there are other ways. Certainly they did not pretend to be sleepy, they were sleepy, and that was a danger. For the moment they popped off, down they fell. The awful thing was that Peter thought this funny. There he goes again.
full caption & clip details
A young adult feminine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as fear, disappointment, awe; style: whispered, narration; average recording, quiet background; genuineness 0.7/6; vocal-burst blend 0.0/10; 20.0s, EN.
EN_B00051_S02556_W000007 · in -19.9 dBFS · gain -0.1 dB · emolia-01235
(distress, sadness, helplessness · measured, very low-energy, no disfluency, whispered) Cried Wendy, looking with horror at the cruel sea far below.
full caption & clip details
A child feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, thin; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; reads as distress, sadness, helplessness; style: whispered, monologue; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 2.3/10; 3.2s, EN.
EN_B00051_S02556_W000008 · in -19.1 dBFS · gain -0.9 dB · emolia-01235
(longing, affection, awe · normal-paced, normally alert, almost no disfluency, narration) Eventually Peter would dive through the air, and catch Michael just before he could strike the sea, and it was lovely the way he did it, but he always waited till the last moment, and he felt it was his cleverness that interested him and not the saving of human life.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as longing, affection, awe; style: narration, monologue; average recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.0/10; 12.8s, EN.
EN_B00051_S02556_W000009 · in -19.1 dBFS · gain -0.9 dB · emolia-01235
(longing, affection, infatuation · normal-paced, normally alert, almost no disfluency, narration) Also he was fond of variety, and the sport that engrossed him one moment would suddenly cease to engage him, so there was always the possibility that the next time you fell he would let you go.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing, affection, infatuation; style: narration, whispered; average recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.0/10; 9.4s, EN.
EN_B00051_S02556_W000010 · in -18.7 dBFS · gain -1.3 dB · emolia-01235
Hope Enthusiasm Optimism(unconstrained axis: Fear)identity −0.03 emotion 81 %   k-B1-k4 · #17

This chain comes from the one-sided rule: only Hope Enthusiasm Optimism had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Hope Enthusiasm Optimism around average — 0.56, higher than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.38.

Nothing was asked of the other axis, and in fact Fear drifts down from 0.91 to 0.37 (-0.54), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.02, then +0.11, then +0.25 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 59 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.659 before conversion and 0.633 after — it fell by 0.027. Neighbour-to-neighbour the worst pair went 0.686 → 0.678. (The earlier render, with segment 1 left raw, scores 0.604 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.379 in the original and +0.309 after conversion — 81 % of the delta retained, which is most of it. On the other named axis, Fear, -0.542 became -0.795.

Quality. Mean predicted overall quality across the segments went 2.88 → 3.17 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.659 → 0.633 -0.027identity cos neighbours 0.686 → 0.678d_b rescored +0.379 → +0.309d_a rescored -0.542 → -0.795d_a mined -0.539d_b mined 0.380min_cos_consec (site) 0.8248min_cos_anchor (site) 0.8712dataset podcastlang enspeaker 326434total 58.2schain gain +3.0 dBseam step 2.3 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, quiet background, moderately variable, some disfluency, average clarity, light breath
(fear · normal-paced, normally alert, slightly relaxed, conversational) (surprised gasp) happen to be (low mumble) around, that's fine. (low mumble) As long as you don't, you know, you're not camping and you're not kick somebody out
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as fear; style: conversational, casual; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 3.8/10; 23.6s, EN.
326434_00153720 · in -21.6 dBFS · gain +1.6 dB · podcast-03632
(contemplation, astonishment surprise · normal-paced, energised, neutral tension, casual) when you're not around or make a stink that you're on vacation and they're sitting at your desk. Because then digital tools will come around eventually with a little bit of that nudge. Like, hey Chris, what you didn't realize today is that someone who is not usually around when you're you around, like is here.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as contemplation, astonishment surprise; style: casual, conversational; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 3.8/10; 18.0s, EN.
326434_00156080 · in -20.9 dBFS · gain +0.8 dB · podcast-06170
(disgust, malevolence malice, contempt · brisk, energised, slightly tense, casual) go have lunch with Phil in this particular place. I'll even make your reservation or whatever it is. Do you
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, fairly guarded; reads as disgust, malevolence malice, contempt; style: casual, conversational; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 1.2/10; 5.8s, EN.
326434_00159952 · in -21.1 dBFS · gain +1.1 dB · podcast-01566
(hope enthusiasm optimism, thankfulness gratitude, concentration · normal-paced, normally alert, neutral tension, casual) Thank you. And just embrace that that, you know, opportunity is there because there's so much of it we can't possibly ingest it all on our own. We need AI to help us point out some of those
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, thankfulness gratitude, concentration; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 4.7/10; 11.4s, EN.
326434_00160680 · in -22.4 dBFS · gain +2.4 dB · podcast-03636
Fear(unconstrained axis: Affection)identity −0.12 emotion 88 %   k-B1-k4 · #18

This chain comes from the one-sided rule: only Fear had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Fear around average — 0.52, higher than 52 % of clips in this corpus — and ends with it strongly present at 0.78, higher than 78 % of clips in this corpus. That is a total rise of 0.27.

Nothing was asked of the other axis, and in fact Affection drifts down from 0.92 to 0.74 (-0.18), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.15, then +0.21, then -0.09 — not a clean run: step 3 moves back the other way by 0.09 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.71 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.79 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.71, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 51 s · es · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.712 before conversion and 0.588 after — it fell by 0.124. Neighbour-to-neighbour the worst pair went 0.726 → 0.664. (The earlier render, with segment 1 left raw, scores 0.665 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.263 in the original and +0.231 after conversion — 88 % of the delta retained, which is most of it. On the other named axis, Affection, -0.327 became +0.797.

Quality. Mean predicted overall quality across the segments went 2.88 → 3.10 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.712 → 0.588 -0.124identity cos neighbours 0.726 → 0.664d_b rescored +0.263 → +0.231d_a rescored -0.327 → +0.797d_a mined -0.177d_b mined 0.268min_cos_consec (site) 0.7918min_cos_anchor (site) 0.7134dataset podcastlang esspeaker 199277total 49.9schain gain +1.7 dBseam step 3.0 dBcrossfades 100/150/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · quiet background, normally alert, moderate pitch range
(affection · normal-paced, neutral tension, moderately variable, casual) esas cosas las hace mucho Lisenda Ardebol, ¿no? Que es una catalana que trabaja en
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as affection; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 5.0/10; 4.8s, ES.
199277_00093920 · in -25.1 dBFS · gain +5.1 dB · podcast-05621
(shame, jealousy and envy, interest · fast, neutral tension, moderately variable, casual) (ahem) (ahem) el mundo.
full caption & clip details
An elderly masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is slightly cool, dark, slightly rough, thin; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as shame, jealousy and envy, interest; style: casual, monologue; below-average recording, quiet background; genuineness 5.9/6; vocal-burst blend 9.2/10; 17.7s, ES.
199277_00094400 · in -23.3 dBFS · gain +3.3 dB · podcast-00727
(disgust, affection, shame · measured, neutral tension, fairly steady, casual) (low mumble) Esa es la (ahem) antropología.
full caption & clip details
An elderly feminine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is slightly cool, slightly dark, slightly rough, thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as disgust, affection, shame; style: casual, monologue; below-average recording, quiet background; genuineness 5.2/6; vocal-burst blend 9.2/10; 19.2s, ES.
199277_00096168 · in -24.5 dBFS · gain +4.5 dB · podcast-00732
(measured, slightly relaxed, fairly steady, monologue) las formas que nosotros tenemos de reconstruir cómo la gente (low mumble) vive y piensa los fenómenos sociales.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 2.4/10; 8.6s, ES.
199277_00098456 · in -25.0 dBFS · gain +5.0 dB · podcast-00732
Hope Enthusiasm Optimism(unconstrained axis: Thankfulness Gratitude)identity +0.33 emotion 98 %   k-B1-k4 · #19

This chain comes from the one-sided rule: only Hope Enthusiasm Optimism had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Hope Enthusiasm Optimism clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.29.

Nothing was asked of the other axis, and in fact Thankfulness Gratitude barely moves at all, sitting near 1.00 throughout.

It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.08, then -0.02 — not a clean run: step 3 moves back the other way by 0.02 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 46 s · ko · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.037 before conversion and 0.368 after — it rose by 0.331. Neighbour-to-neighbour the worst pair went 0.054 → 0.385. (The earlier render, with segment 1 left raw, scores 0.246 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.288 in the original and +0.282 after conversion — 98 % of the delta retained, which is essentially all of it. On the other named axis, Thankfulness Gratitude, -0.028 became -0.042.

Quality. Mean predicted overall quality across the segments went 2.88 → 3.10 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.037 → 0.368 +0.331identity cos neighbours 0.054 → 0.385d_b rescored +0.288 → +0.282d_a rescored -0.028 → -0.042d_a mined -0.028d_b mined 0.288min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang kospeaker KO_SqPNmyGJAcktotal 44.6schain gain +1.9 dBseam step 1.0 dBcrossfades 150/150/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, slightly bright, fairly smooth
(thankfulness gratitude, affection, contentment · normal-paced, normally alert, slightly relaxed, casual) Anyway, (ahem) uh, thank you for listening to (ahem) this talk.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as thankfulness gratitude, affection, contentment; style: casual, conversational; good recording, no background noise; genuineness 2.4/6; vocal-burst blend 1.8/10; 4.3s, KO.
KO_SqPNmyGJAck_W000140 · in -20.0 dBFS · gain -0.0 dB · emolia-03264
(pleasure ecstasy, contentment, infatuation · brisk, normally alert, neutral tension, casual) 카렌 샌들어님의 발표를 영상으로 함께 만나보셨습니다. (surprised gasp) 어, (ahem) 지금 카렌 샌들어님이 온라인으로 접속을 하고 계신대요. 자, 만나뵙기 전에 여러분 혹시 현장에서 또 질문하실 분이 혹시 계실까요? 손으로 한번 들어주시겠어요? 아, 네. 알겠습니다.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as pleasure ecstasy, contentment, infatuation; style: casual, dramatic; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 5.9/10; 16.3s, KO.
KO_SqPNmyGJAck_W000141 · in -23.9 dBFS · gain +3.9 dB · emolia-03264
(pleasure ecstasy, elation, hope enthusiasm optimism · brisk, energised, neutral tension, casual) (wistful sigh) I can hear you. (ahem) Hello.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as pleasure ecstasy, elation, hope enthusiasm optimism; style: casual, dramatic; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 4.8/10; 16.7s, KO.
KO_SqPNmyGJAck_W000142 · in -20.7 dBFS · gain +0.7 dB · emolia-03264
(hope enthusiasm optimism, thankfulness gratitude, sexual lust · brisk, energised, neutral tension, casual) 자 그러면 발표 내용을 토대로 저희가 질문을 좀 드리도록 하겠습니다. 지금부터 질문하는 내용의 답변을 좀 부탁드리겠습니다.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism, thankfulness gratitude, sexual lust; style: casual, playful; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 7.2/10; 8.0s, KO.
KO_SqPNmyGJAck_W000144 · in -22.1 dBFS · gain +2.1 dB · emolia-03264
Fear(unconstrained axis: Infatuation)identity −0.08 emotion 69 %   k-B1-k4 · #20

This chain comes from the one-sided rule: only Fear had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Fear around average — 0.45, lower than 55 % of clips in this corpus — and ends with it clearly present at 0.75, higher than 75 % of clips in this corpus. That is a total rise of 0.30.

Nothing was asked of the other axis, and in fact Infatuation drifts down from 0.94 to 0.48 (-0.46), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.22, then -0.08, then +0.17 — not a clean run: step 2 moves back the other way by 0.08 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 43 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.868 before conversion and 0.791 after — it fell by 0.076. Neighbour-to-neighbour the worst pair went 0.940 → 0.868. (The earlier render, with segment 1 left raw, scores 0.739 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.303 in the original and +0.208 after conversion — 69 % of the delta retained. On the other named axis, Infatuation, -0.463 became -0.390.

Quality. Mean predicted overall quality across the segments went 3.03 → 3.13 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.868 → 0.791 -0.076identity cos neighbours 0.940 → 0.868d_b rescored +0.303 → +0.208d_a rescored -0.463 → -0.390d_a mined -0.463d_b mined 0.303min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_sx0RUIeNuZMtotal 41.9schain gain +1.3 dBseam step 0.8 dBcrossfades 100/100/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed, no disfluency
(infatuation · normal-paced, steady, no audible breath, newsreading) By the time the first European settlers arrived, the people – subsequently called the Duwamish Tribe – occupied at least 17 villages in the areas around Elliott Bay.The first European to visit the Seattle area was George Vancouver, in May 1792 during his 1791–95 expedition to chart the Pacific Northwest.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, no audible breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 19.9s, EN.
EN_sx0RUIeNuZM_W000022 · in -15.6 dBFS · gain -4.4 dB · emolia-01406
(normal-paced, fairly steady, light breath, formal) In 1851, a large party led by Luther Collins made a location on land at the mouth of the Duwamish River, they formally claimed it on September 14, 1851
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.4s, EN.
EN_sx0RUIeNuZM_W000023 · in -14.0 dBFS · gain -6.0 dB · emolia-01406
(emotional numbness · normal-paced, fairly steady, light breath, formal) 13 days later, members of the Collins Party on the way to their claim passed three scouts of the Denny Party
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 6.3s, EN.
EN_sx0RUIeNuZM_W000024 · in -15.0 dBFS · gain -5.0 dB · emolia-01406
(measured, steady, light breath, formal) Members of the Denny Party claimed land on Alki Point on September 28, 1851
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, casual; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.2/10; 5.8s, EN.
EN_sx0RUIeNuZM_W000025 · in -15.0 dBFS · gain -5.0 dB · emolia-01406