emotion__B1__T0.25__C0.25__INTERNAL — voice-corrected

Manifest tier. emotion, rule B1, T=0.25, step cap 0.25. Population 9,393,332 chains (106,540 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 7,358,803.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_emotion__B1__T0.25__C0.25__INTERNAL.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
50segments re-voiced
0.763 → 0.781median worst-to-anchor identity cosine
113 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Contentment(unconstrained axis: Triumph)identity +0.19 emotion 19 %   emotion__B1__T0.25__C0.25__INTERNAL · #1

This chain comes from the one-sided rule: only Contentment had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Contentment around average — 0.48, lower than 52 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.45.

Nothing was asked of the other axis, and in fact Triumph barely moves at all, sitting near 0.89 throughout.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.87 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.87 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 27 s · ko · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.669 before conversion and 0.864 after — it rose by 0.195. Neighbour-to-neighbour the worst pair went 0.601 → 0.803. (The earlier render, with segment 1 left raw, scores 0.604 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.456 in the original and +0.087 after conversion — 19 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Triumph, +0.027 became +0.742.

Quality. Mean predicted overall quality across the segments went 2.84 → 3.19 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.669 → 0.864 +0.195identity cos neighbours 0.601 → 0.803d_b rescored +0.456 → +0.087d_a rescored +0.027 → +0.742d_a mined 0.027d_b mined 0.455min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang kospeaker KO_t7ppwjsf6Dktotal 26.3schain gain +1.0 dBseam step 0.7 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, measured, normally alert, slightly relaxed
(steady, little disfluency, fairly narrow pitch, monologue) 이미 언론상에 예고됐던 내용 외에 추가로 규제 완화가 되는 부분까지 자세히 알아보겠습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, little disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.0/10; 8.5s, KO.
KO_t7ppwjsf6Dk_W000001 · in -20.9 dBFS · gain +0.9 dB · emolia-03204
(longing, pride, triumph · fairly steady, frequent disfluency, moderate pitch range, monologue) 네, 안녕하세요. 박찬웅의 부동산 쇼, 우잔남 박찬웅입니다.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as longing, pride, triumph; style: monologue, conversational; average recording, no background noise; genuineness 4.1/6; vocal-burst blend 2.4/10; 5.3s, KO.
KO_t7ppwjsf6Dk_W000002 · in -17.3 dBFS · gain -2.7 dB · emolia-03204
(contentment, triumph · fairly steady, frequent disfluency, moderate pitch range, monologue) 지난 12월 23일이었습니다. 음, (low mumble) 각종 세금 규제 완화 내용을 담고 있는 세금별 개정법안이 국회 본회의를 통과했습니다. 이제 확정이 된 거죠.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contentment, triumph; style: monologue, didactic; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 2.8/10; 13.0s, KO.
KO_t7ppwjsf6Dk_W000003 · in -21.3 dBFS · gain +1.3 dB · emolia-03204
Concentration(unconstrained axis: Fear)identity −0.02 emotion 169 %   emotion__B1__T0.25__C0.25__INTERNAL · #2

This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Concentration clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.35.

Nothing was asked of the other axis, and in fact Fear drifts down from 0.97 to 0.88 (-0.08), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.21, then +0.06, then +0.08 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 83 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.863 before conversion and 0.845 after — it fell by 0.018. Neighbour-to-neighbour the worst pair went 0.818 → 0.788. (The earlier render, with segment 1 left raw, scores 0.757 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.348 in the original and +0.587 after conversion — 169 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Fear, -0.082 became -0.092.

Quality. Mean predicted overall quality across the segments went 3.03 → 3.25 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.863 → 0.845 -0.018identity cos neighbours 0.818 → 0.788d_b rescored +0.348 → +0.587d_a rescored -0.082 → -0.092d_a mined -0.082d_b mined 0.349min_cos_consec (site) 0.8757min_cos_anchor (site) 0.8757dataset podcastlang enspeaker 815568total 81.8schain gain +4.6 dBseam step 1.2 dBcrossfades 150/150/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, slightly relaxed, moderate pitch range, light breath
(fear, helplessness, interest · normal-paced, subdued, fairly steady, casual) We have nothing that can do something like that. Recently, there have now been drones made and talked about openly, and actually these are US (low mumble) uh drones, just shown on a (low mumble) um I saw in a on a military video recently, where they can go drones can be underwater travel and then come out of the water and go do the attack.
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as fear, helplessness, interest; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 9.5/10; 21.3s, EN.
815568_00491080 · in -18.6 dBFS · gain -1.4 dB · podcast-01853
(contemplation, longing, doubt · normal-paced, normally alert, moderately variable, conversational) But that's only been developed in the last you know a few years, not something from 50, 60 years ago. So those are the kinds of things that people see. And again, it's it's you know, you you ask me what I
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contemplation, longing, doubt; style: conversational, didactic; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 3.1/10; 13.7s, EN.
815568_00493208 · in -19.1 dBFS · gain -0.9 dB · podcast-01837
(contemplation, awe, malevolence malice · measured, subdued, fairly steady, monologue) Think of as real. That those anecdotes are to me stories. And why I get interested in the medical or the material side is it's something I can repeat. I can't repeat these pilot observations, but I can repeat experiments on materials or experiments on not experiments on human, but reading the humans who've been harmed. Now, the thing about Skywatcher is (low mumble)
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as contemplation, awe, malevolence malice; style: monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 5.5/10; 29.9s, EN.
815568_00494576 · in -18.6 dBFS · gain -1.4 dB · podcast-06146
(concentration, contemplation, relief · measured, subdued, fairly steady, whispered) that we (low mumble) uh, at least in a limited sense, have a signal that can be released that sometimes it seems to attract these objects. And so that's where the repeatability attempt is coming in.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as concentration, contemplation, relief; style: whispered, monologue; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 1.7/10; 17.4s, EN.
815568_00497560 · in -20.1 dBFS · gain +0.1 dB · podcast-05411
Doubt(unconstrained axis: Anger)identity −0.03 emotion 83 %   emotion__B1__T0.25__C0.25__INTERNAL · #3

This chain comes from the one-sided rule: only Doubt had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Doubt clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.29.

Nothing was asked of the other axis, and in fact Anger drifts down from 0.93 to 0.64 (-0.29), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then -0.07, then +0.12 — not a clean run: step 2 moves back the other way by 0.07 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 41 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.775 before conversion and 0.746 after — it fell by 0.029. Neighbour-to-neighbour the worst pair went 0.775 → 0.697. (The earlier render, with segment 1 left raw, scores 0.533 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.289 in the original and +0.240 after conversion — 83 % of the delta retained, which is most of it. On the other named axis, Anger, -0.289 became -0.308.

Quality. Mean predicted overall quality across the segments went 3.01 → 3.14 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.775 → 0.746 -0.029identity cos neighbours 0.775 → 0.697d_b rescored +0.289 → +0.240d_a rescored -0.289 → -0.308d_a mined -0.289d_b mined 0.288min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00009_S06590total 40.2schain gain +0.4 dBseam step 4.2 dBcrossfades 150/100/100 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, fairly steady
(anger · slightly relaxed, casual, conversational) Uh, (low mumble) lower in, more lower to moderate income people who probably can't get a house, you know, they're, they're, they're renters for life, not necessarily by choice. And then D is, uh, (ahem) if you go, you better be packing heat.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as anger; style: casual, conversational; good recording, quiet background; genuineness 2.6/6; vocal-burst blend 3.3/10; 12.5s, EN.
EN_B00009_S06590_W000026 · in -17.9 dBFS · gain -2.1 dB · emolia-00427
(embarrassment, teasing, infatuation · neutral tension, casual, conversational) Generally not, (low mumble) uhm, I mean, some people, we joke about it occasionally, but yeaah, D and anything below C, just stay away from.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as embarrassment, teasing, infatuation; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 4.5/10; 7.6s, EN.
EN_B00009_S06590_W000027 · in -19.1 dBFS · gain -0.9 dB · emolia-00427
(slightly relaxed, casual, conversational) (low mumble) Uh, self-managed is a great one. That, that (ahem) we absolutely, I mean, when you go from self-managed
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 1.9/10; 7.6s, EN.
EN_B00009_S06590_W000028 · in -18.8 dBFS · gain -1.2 dB · emolia-00427
(doubt · slightly relaxed, conversational, casual) Well, I shouldn't say that because I don't want to discount there, there are some professional operators out there that, that manage their own property. So when I say self-manage, I mean someone that maybe owns one property or two and they think they can save money.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as doubt; style: conversational, casual; good recording, quiet background; genuineness 4.3/6; vocal-burst blend 6.9/10; 12.9s, EN.
EN_B00009_S06590_W000029 · in -18.7 dBFS · gain -1.3 dB · emolia-00427
Hope Enthusiasm Optimism(unconstrained axis: Infatuation)identity +0.05 emotion 31 %   emotion__B1__T0.25__C0.25__INTERNAL · #4

This chain comes from the one-sided rule: only Hope Enthusiasm Optimism had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Hope Enthusiasm Optimism clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.25.

Nothing was asked of the other axis, and in fact Infatuation drifts down from 0.98 to 0.89 (-0.09), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.01, then +0.25 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.82 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.82 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 19 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.699 before conversion and 0.744 after — it rose by 0.045. Neighbour-to-neighbour the worst pair went 0.721 → 0.744. (The earlier render, with segment 1 left raw, scores 0.623 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.250 in the original and +0.077 after conversion — 31 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Infatuation, -0.096 became -0.136.

Quality. Mean predicted overall quality across the segments went 2.28 → 2.88 (+0.60) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.699 → 0.744 +0.045identity cos neighbours 0.721 → 0.744d_b rescored +0.250 → +0.077d_a rescored -0.096 → -0.136d_a mined -0.091d_b mined 0.250min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN__xJegruF_UEtotal 18.4schain gain +1.9 dBseam step 0.8 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · slightly thin, average recording, quiet background, wide pitch range
(infatuation, jealousy and envy, affection · measured, very low-energy, relaxed, whispered) See, I'm telling you. And see how much better it looks though? It's clean.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is slightly warm, neutral-bright, smooth, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, audible breath; affect is mildly positive, submissive, slightly vulnerable; reads as infatuation, jealousy and envy, affection; style: whispered, casual; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 3.9/10; 4.4s, EN.
EN__xJegruF_UE_W000106 · in -21.3 dBFS · gain +1.3 dB · emolia-01157
(fatigue exhaustion, infatuation, sexual lust · normal-paced, very low-energy, relaxed, casual) So my face is all waxed and ready to go. I'm gonna do a mask right now because I know my skin gets really sensitive and like, I don't know if you can see but it's really really red up here so.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as fatigue exhaustion, infatuation, sexual lust; style: casual, conversational; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 5.0/10; 10.1s, EN.
EN__xJegruF_UE_W000107 · in -22.0 dBFS · gain +2.0 dB · emolia-01157
(normal-paced, normally alert, slightly relaxed, casual) I'm gonna do this mask from Dr. Jart and it is the brightening solution.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 2.7/10; 4.4s, EN.
EN__xJegruF_UE_W000108 · in -19.8 dBFS · gain -0.2 dB · emolia-01157
Concentration(unconstrained axis: Helplessness)identity −0.09 emotion 87 %   emotion__B1__T0.25__C0.25__INTERNAL · #5

This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Concentration clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.38.

Nothing was asked of the other axis, and in fact Helplessness drifts down from 0.93 to 0.41 (-0.52), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.19, then +0.14, then +0.04, then +0.01 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.57 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.57 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 54 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.612 before conversion and 0.522 after — it fell by 0.090. Neighbour-to-neighbour the worst pair went 0.616 → 0.576. (The earlier render, with segment 1 left raw, scores 0.514 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.385 in the original and +0.337 after conversion — 87 % of the delta retained, which is most of it. On the other named axis, Helplessness, -0.518 became -0.353.

Quality. Mean predicted overall quality across the segments went 2.69 → 2.99 (+0.31) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.612 → 0.522 -0.090identity cos neighbours 0.616 → 0.576d_b rescored +0.385 → +0.337d_a rescored -0.518 → -0.353d_a mined -0.518d_b mined 0.385min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00019_S05302total 52.8schain gain +2.9 dBseam step 3.2 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · neutral-toned, fairly smooth, slightly relaxed, frequent disfluency
(helplessness · slow, subdued, steady, monologue) Of the ball, centred at X and of radius, epsilon.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness; style: monologue, formal; average recording, no background noise; genuineness 2.2/6; vocal-burst blend 1.0/10; 4.4s, EN.
EN_B00019_S05302_W000164 · in -24.0 dBFS · gain +4.0 dB · emolia-00614
(impatience and irritability, anger, emotional numbness · measured, normally alert, fairly steady, didactic) Right, so the consequence of all this argument is this fact.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as impatience and irritability, anger, emotional numbness; style: didactic, monologue; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 0.0/10; 4.8s, EN.
EN_B00019_S05302_W000165 · in -19.1 dBFS · gain -0.9 dB · emolia-00614
(concentration, pride, contemplation · measured, normally alert, fairly steady, didactic) And here what we did is simply we played with the linearity of m to replace r by r epsilon and 1 by epsilon and then we translate everything by y.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, pride, contemplation; style: didactic, monologue; good recording, quiet background; genuineness 2.4/6; vocal-burst blend 0.7/10; 10.7s, EN.
EN_B00019_S05302_W000166 · in -21.3 dBFS · gain +1.3 dB · emolia-00614
(concentration, disappointment, triumph · measured, normally alert, moderately variable, didactic) We reach this conclusion, the proof it's finished, because, well, we know that the ball of radius epsilon at x is contained in G, so the image of this ball is contained in the image of mg.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, slightly dark, fairly smooth, thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, disappointment, triumph; style: didactic, monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.0/10; 13.3s, EN.
EN_B00019_S05302_W000167 · in -19.0 dBFS · gain -1.0 dB · emolia-00614
(concentration, triumph · measured, normally alert, fairly steady, didactic) Center dot y and of radius r epsilon, it's contained in mg, and this is exactly what we wanted to prove, right? We wanted to prove that there exists a strictly positive s, such that the ball center dot y and of radius s, it's contained in mg, we proved that with s equal to r epsilon, r
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration, triumph; style: didactic, monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.9/10; 20.3s, EN.
EN_B00019_S05302_W000168 · in -19.4 dBFS · gain -0.6 dB · emolia-00614
Pride(unconstrained axis: Embarrassment)identity −0.02 emotion 202 %   emotion__B1__T0.25__C0.25__INTERNAL · #6

This chain comes from the one-sided rule: only Pride had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Pride clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.36.

Nothing was asked of the other axis, and in fact Embarrassment drifts down from 0.96 to 0.84 (-0.12), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.14 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 50 s · hr · eurospeech

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.923 before conversion and 0.904 after — it fell by 0.019. Neighbour-to-neighbour the worst pair went 0.926 → 0.922. (The earlier render, with segment 1 left raw, scores 0.784 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.358 in the original and +0.724 after conversion — 202 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Embarrassment, -0.124 became +0.013.

Quality. Mean predicted overall quality across the segments went 3.00 → 3.29 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.923 → 0.904 -0.019identity cos neighbours 0.926 → 0.922d_b rescored +0.358 → +0.724d_a rescored -0.124 → +0.013d_a mined -0.124d_b mined 0.358min_cos_consec (site) —min_cos_anchor (site) —dataset eurospeechlang hrspeaker croatia_20240709091640-126total 49.6schain gain +1.5 dBseam step 3.2 dBcrossfades 150/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, average recording, quiet background, normally alert, neutral tension, moderate pitch range, light breath
(embarrassment, shame · measured, fairly steady, frequent disfluency, casual) (low mumble) još jednu stvar ću spomenuti (ahem) u ovo malo vremena koliko imam, u Dubrovniku npr. kvaliteta zraka vezana uz (ahem) (low mumble) brodove na kružnim putovanjima uz kruzere, (low mumble) bio je (low mumble) reakcija dakle (low mumble)
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as embarrassment, shame; style: casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 4.1/10; 18.5s, HR.
croatia_20240709091640-12681_11239648_11258112 · in -15.7 dBFS · gain -4.3 dB · eurospeech-01623
(embarrassment, shame · normal-paced, moderately variable, some disfluency, casual) odnosno primjedbe građana dakle da i kruzeri (low mumble) pridonose onečišćenju zraka. (low mumble) Međutim uz neke (low mumble) rigorozne mjere navodno (low mumble) su oni popravili zapravo, rjeđe dolaze u grad ali imaju i procedure unaprijeđene, međutim i dalje
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as embarrassment, shame; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 9.7/10; 16.1s, HR.
croatia_20240709091640-12681_11258112_11274192 · in -16.1 dBFS · gain -3.9 dB · eurospeech-01623
(pride, thankfulness gratitude, triumph · normal-paced, fairly steady, some disfluency, casual) građani protestiraju i tuže se zapravo na kvalitetu zraka (ahem) izričito uzrokovanom, oni misle (low mumble) kruzerima jel. Dakle i to je poziv da se provjeri stanje kvalitete zraka kao i u Dubrovniku. Hvala.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, thankfulness gratitude, triumph; style: casual, monologue; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 6.0/10; 15.4s, HR.
croatia_20240709091640-12681_11274192_11289616 · in -19.6 dBFS · gain -0.4 dB · eurospeech-01623
Confusion(unconstrained axis: Disappointment)identity −0.10 emotion 97 %   emotion__B1__T0.25__C0.25__INTERNAL · #7

This chain comes from the one-sided rule: only Confusion had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Confusion clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.38.

Nothing was asked of the other axis, and in fact Disappointment drifts down from 1.00 to 0.39 (-0.60), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.17 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.74 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.67 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.74, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 39 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.555 before conversion and 0.458 after — it fell by 0.096. Neighbour-to-neighbour the worst pair went 0.589 → 0.544. (The earlier render, with segment 1 left raw, scores 0.417 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.378 in the original and +0.367 after conversion — 97 % of the delta retained, which is essentially all of it. On the other named axis, Disappointment, -0.603 became -0.572.

Quality. Mean predicted overall quality across the segments went 2.91 → 3.01 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.555 → 0.458 -0.096identity cos neighbours 0.589 → 0.544d_b rescored +0.378 → +0.367d_a rescored -0.603 → -0.572d_a mined -0.603d_b mined 0.378min_cos_consec (site) 0.6668min_cos_anchor (site) 0.7405dataset podcastlang enspeaker 876700total 38.3schain gain +2.3 dBseam step 1.4 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, average recording, moderately variable
(disappointment, jealousy and envy, fatigue exhaustion · measured, very low-energy, relaxed, casual) the the clip was broken. Fast forward to a few months later. You asked me to do you the favor and drive you to your dentist's office. And I'm like, alright, cool. You know, it's in the city, it's easier to maneuver on a motorcycle than it is where you're big old Durango. (low mumble) Um but I can't wear my I don't like wearing my hat under my helmet. So I gave you my hat and I'm like, put it in your book bag. You can (low mumble) uh put it in your book bag and (low mumble) um
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is neutral-toned, dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as disappointment, jealousy and envy, fatigue exhaustion; style: casual, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 6.5/10; 27.8s, EN.
876700_00167408 · in -34.9 dBFS · gain +14.8 dB · podcast-03502
(embarrassment · slow, very low-energy, relaxed, monologue) She didn't. She wanted to (low mumble) keep us secure, so she tried to strap it.
full caption & clip details
A child somewhat masculine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, frequent disfluency, narrow pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as embarrassment; style: monologue, casual; average recording, no background noise; genuineness 2.3/6; vocal-burst blend 0.7/10; 6.8s, EN.
876700_00170832 · in -36.2 dBFS · gain +16.2 dB · podcast-03448
(confusion, sourness, amusement · normal-paced, normally alert, slightly relaxed, casual) I was not aware. Well, she was aware, she just forgot. Because it's been a few months.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as confusion, sourness, amusement; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 2.7/10; 4.0s, EN.
876700_00172128 · in -32.8 dBFS · gain +12.8 dB · podcast-03449
Confusion(unconstrained axis: Infatuation)identity +0.02 emotion 132 %   emotion__B1__T0.25__C0.25__INTERNAL · #8

This chain comes from the one-sided rule: only Confusion had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Confusion around average — 0.49, lower than 51 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.37.

Nothing was asked of the other axis, and in fact Infatuation barely moves at all, sitting near 0.81 throughout.

It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.24 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.95 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.95 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 29 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.837 before conversion and 0.862 after — it rose by 0.025. Neighbour-to-neighbour the worst pair went 0.837 → 0.839. (The earlier render, with segment 1 left raw, scores 0.796 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.370 in the original and +0.487 after conversion — 132 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Infatuation, +0.030 became +0.124.

Quality. Mean predicted overall quality across the segments went 3.19 → 3.34 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.837 → 0.862 +0.025identity cos neighbours 0.837 → 0.839d_b rescored +0.370 → +0.487d_a rescored +0.030 → +0.124d_a mined 0.030d_b mined 0.370min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00062_S06616total 28.5schain gain -0.2 dBseam step 0.8 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, no background noise, normally alert, slightly relaxed, fairly steady, no disfluency
(normal-paced, moderate pitch range, formal, monologue) 荣夫人和柔夫人是最高兴的一别数年,他们几个好姐妹终于能得以重聚了。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 2.0/10; 6.6s, ZH.
ZH_B00062_S06616_W000012 · in -17.0 dBFS · gain -3.0 dB · emolia-03898
(triumph, concentration · normal-paced, narrow pitch range, formal, monologue) 洛青莲有了身孕,这次他非常谨慎,决定要将此事保密。就连喝安胎药的事情都万般谨慎,他担心之前的事情再次发生。所以这件事只有身边的下人,还有柔夫人荣夫人之情。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, narrow pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, concentration; style: formal, monologue; average recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.1/10; 15.1s, ZH.
ZH_B00062_S06616_W000013 · in -17.9 dBFS · gain -2.1 dB · emolia-03898
(measured, moderate pitch range, narration, cartoonish) 只有羽翼丰满才能保住孩子,所以就算拼了性命,他也要保住自己的第二个孩子。
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, cartoonish; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 1.5/10; 7.2s, ZH.
ZH_B00062_S06616_W000014 · in -18.6 dBFS · gain -1.4 dB · emolia-03898
Thankfulness Gratitude(unconstrained axis: Infatuation)identity −0.02 emotion 100 %   emotion__B1__T0.25__C0.25__INTERNAL · #9

This chain comes from the one-sided rule: only Thankfulness Gratitude had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Thankfulness Gratitude clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 92 % of clips in this corpus. That is a total rise of 0.29.

Nothing was asked of the other axis, and in fact Infatuation drifts down from 0.92 to 0.84 (-0.08), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.07 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.94 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.94 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 31 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.887 before conversion and 0.868 after — it fell by 0.019. Neighbour-to-neighbour the worst pair went 0.855 → 0.834. (The earlier render, with segment 1 left raw, scores 0.853 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.285 in the original and +0.284 after conversion — 100 % of the delta retained, which is essentially all of it. On the other named axis, Infatuation, -0.078 became -0.007.

Quality. Mean predicted overall quality across the segments went 2.93 → 3.18 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.887 → 0.868 -0.019identity cos neighbours 0.855 → 0.834d_b rescored +0.285 → +0.284d_a rescored -0.078 → -0.007d_a mined -0.078d_b mined 0.285min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00065_S04382total 30.0schain gain +0.8 dBseam step 1.6 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, no disfluency
(infatuation · measured, steady, moderate pitch range, formal) 第十六集常毅和空明联手报仇林浩,清没有救出纪云河,还被明清打伤。他决定回万花谷。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.5/10; 7.8s, ZH.
ZH_B00065_S04382_W000000 · in -19.1 dBFS · gain -0.9 dB · emolia-03924
(infatuation · fast, fairly steady, narrow pitch range, formal) 洛锦桑坚决不干,想留在露台山,继续营救纪云河。林浩清只好承认纪云河根本不跟他走,还让他好好照顾洛锦桑。骆锦桑只好作罢。林浩清心里暗暗发誓,把纪云河安然无恙,从这里带走。
full caption & clip details
A young adult feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, narrow pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation; style: formal, monologue; average recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.0/10; 15.4s, ZH.
ZH_B00065_S04382_W000001 · in -19.4 dBFS · gain -0.6 dB · emolia-03924
(thankfulness gratitude · normal-paced, fairly steady, moderate pitch range, formal) 尽管明清下令隐瞒纪云河被扎的消息,林浩清闯先师府救人的事很快传到蜀灵耳朵里。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.0/10; 7.1s, ZH.
ZH_B00065_S04382_W000002 · in -19.3 dBFS · gain -0.7 dB · emolia-03924
Interest(unconstrained axis: Pain)identity +0.30 emotion 119 %   emotion__B1__T0.25__C0.25__INTERNAL · #10

This chain comes from the one-sided rule: only Interest had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Interest clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.26.

Nothing was asked of the other axis, and in fact Pain drifts down from 0.89 to 0.07 (-0.82), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.21, then -0.01, then +0.06, then +0.00 — not a clean run: step 2 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.11 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.11 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 61 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.245 before conversion and 0.542 after — it rose by 0.297. Neighbour-to-neighbour the worst pair went 0.114 → 0.528. (The earlier render, with segment 1 left raw, scores 0.521 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.261 in the original and +0.312 after conversion — 119 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Pain, -0.819 became -0.810.

Quality. Mean predicted overall quality across the segments went 2.78 → 3.14 (+0.36) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.245 → 0.542 +0.297identity cos neighbours 0.114 → 0.528d_b rescored +0.261 → +0.312d_a rescored -0.819 → -0.810d_a mined -0.819d_b mined 0.261min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_HlCQ5YZdCFQtotal 60.0schain gain +3.9 dBseam step 1.2 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, good recording, normally alert, slightly relaxed, fairly steady, some disfluency
(normal-paced, average clarity, monologue, casual) (ahem) Uhm, so you can see from a, (ahem) like, (low mumble) uh, compliance or regulatory (ahem) point perspective, you know where the metadata is stored in Azure.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; good recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.2/10; 7.1s, EN.
EN_HlCQ5YZdCFQ_W000033 · in -16.2 dBFS · gain -3.8 dB · emolia-01965
(concentration · normal-paced, average clarity, casual, monologue) Physical location is new, specifically for the on-prem servers. This allows customers to tag the servers or, (low mumble) uhm, specifically indicate which data center they are in. (ahem) Uh, so if there's something, this is really about ease of management.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration; style: casual, monologue; good recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.1/10; 13.7s, EN.
EN_HlCQ5YZdCFQ_W000034 · in -17.8 dBFS · gain -2.2 dB · emolia-01965
(brisk, average clarity, casual, monologue) Okay, that's pretty cool. So customers could not just add a date name over their data center. So they could even, like, for example, also add a room of the location or even the rack name or the rack number (ahem) for their server.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 2.6/10; 11.4s, EN.
EN_HlCQ5YZdCFQ_W000035 · in -13.9 dBFS · gain -6.1 dB · emolia-01965
(interest, hope enthusiasm optimism, concentration · brisk, clear, casual, conversational) Yeah, absolutely. So this is really for the customer to easily identify where that resource is. If something happens to that server, they can go. If they need to physically access, they know exactly where they need to be.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as interest, hope enthusiasm optimism, concentration; style: casual, conversational; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 2.7/10; 11.0s, EN.
EN_HlCQ5YZdCFQ_W000036 · in -17.0 dBFS · gain -3.0 dB · emolia-01965
(interest, relief, hope enthusiasm optimism · normal-paced, average clarity, monologue, didactic) Here we also (ahem) allow customers to choose the operating systems as I didn't really specifically spell it out, but (ahem) as always in Azure, we are trying to embrace Windows as well as Linux. Same for here that (ahem) we build two packages (low mumble) for agents to onboard either Windows Server or Linux Server.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, relief, hope enthusiasm optimism; style: monologue, didactic; good recording, quiet background; genuineness 1.7/6; vocal-burst blend 1.2/10; 17.6s, EN.
EN_HlCQ5YZdCFQ_W000037 · in -16.9 dBFS · gain -3.0 dB · emolia-01965
Emotional Numbness(unconstrained axis: Fatigue Exhaustion)identity −0.13 emotion 54 %   emotion__B1__T0.25__C0.25__INTERNAL · #11

This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Emotional Numbness clearly present — 0.59, higher than 59 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.39.

Nothing was asked of the other axis, and in fact Fatigue Exhaustion drifts down from 0.81 to 0.50 (-0.32), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.07, then +0.17, then +0.15 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.91 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.91 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 23 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.789 before conversion and 0.660 after — it fell by 0.129. Neighbour-to-neighbour the worst pair went 0.729 → 0.747. (The earlier render, with segment 1 left raw, scores 0.578 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.387 in the original and +0.209 after conversion — 54 % of the delta retained. On the other named axis, Fatigue Exhaustion, -0.317 became -0.580.

Quality. Mean predicted overall quality across the segments went 2.68 → 2.92 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.789 → 0.660 -0.129identity cos neighbours 0.729 → 0.747d_b rescored +0.387 → +0.209d_a rescored -0.317 → -0.580d_a mined -0.317d_b mined 0.387min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_OXBH-8vMCRYtotal 21.8schain gain +1.2 dBseam step 1.1 dBcrossfades 100/150/100 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normally alert, slightly relaxed
(measured, frequent disfluency, somewhat unclear, monologue) (ahem) Uh, here at the September 13th city council meeting, uh, (low mumble) subsequently we were informed there was a. (low mumble)
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 0.5/10; 6.4s, EN.
EN_OXBH-8vMCRY_W000067 · in -18.1 dBFS · gain -1.9 dB · emolia-00872
(normal-paced, some disfluency, average clarity, casual) An issue with the legal notices that were mailed out. Accordingly we're back here tonight.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, authoritative; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.4/10; 4.9s, EN.
EN_OXBH-8vMCRY_W000068 · in -17.4 dBFS · gain -2.5 dB · emolia-00872
(normal-paced, some disfluency, average clarity, casual) (low mumble) So I don't have (low mumble) the rest of the team that was here that night.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; style: casual, conversational; average recording, quiet background; no dominant emotion; genuineness 3.2/6; vocal-burst blend 2.3/10; 3.3s, EN.
EN_OXBH-8vMCRY_W000069 · in -16.9 dBFS · gain -3.1 dB · emolia-00872
(emotional numbness, pride, triumph · measured, some disfluency, average clarity, monologue) Including but not limited to co-counsel on the matter, Mr. Jim Ward of Nutter. (ahem) Uh, the locus is 840 Winter Street.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as emotional numbness, pride, triumph; style: monologue, authoritative; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.0/10; 7.6s, EN.
EN_OXBH-8vMCRY_W000070 · in -21.1 dBFS · gain +1.1 dB · emolia-00872
Fear(unconstrained axis: Fatigue Exhaustion)identity −0.04 emotion 124 %   emotion__B1__T0.25__C0.25__INTERNAL · #12

This chain comes from the one-sided rule: only Fear had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Fear around average — 0.49, right about the corpus median — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.36.

Nothing was asked of the other axis, and in fact Fatigue Exhaustion barely moves at all, sitting near 0.80 throughout.

It takes 3 clips to get there. Clip to clip the moves are +0.14, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.92 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.92 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 14 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.915 before conversion and 0.875 after — it fell by 0.040. Neighbour-to-neighbour the worst pair went 0.908 → 0.869. (The earlier render, with segment 1 left raw, scores 0.846 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.356 in the original and +0.442 after conversion — 124 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Fatigue Exhaustion, +0.048 became +0.213.

Quality. Mean predicted overall quality across the segments went 2.94 → 3.01 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.915 → 0.875 -0.040identity cos neighbours 0.908 → 0.869d_b rescored +0.356 → +0.442d_a rescored +0.048 → +0.213d_a mined 0.048d_b mined 0.356min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00010_S08116total 13.8schain gain +1.8 dBseam step 0.1 dBcrossfades 100/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady
(normal-paced, no disfluency, clear, formal) 可是,不管舆论如何,刘平只希望能把问题解决掉。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 3.3/10; 3.9s, ZH.
ZH_B00010_S08116_W000002 · in -18.4 dBFS · gain -1.6 dB · emolia-03378
(normal-paced, some disfluency, average clarity, casual) 刘平的法庭申诉也开启了华为历史上的一个第一。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, playful; good recording, no background noise; genuineness 3.2/6; vocal-burst blend 2.4/10; 4.1s, ZH.
ZH_B00010_S08116_W000003 · in -17.8 dBFS · gain -2.2 dB · emolia-03378
(measured, no disfluency, clear, authoritative) 二零零三年五月二十七日,华为遭遇了有史以来第一起与股权争执有关的案件。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, didactic; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 2.2/10; 6.2s, ZH.
ZH_B00010_S08116_W000004 · in -17.9 dBFS · gain -2.1 dB · emolia-03378
Pleasure Ecstasy(unconstrained axis: Embarrassment)identity +0.55 emotion 251 %   emotion__B1__T0.25__C0.25__INTERNAL · #13

This chain comes from the one-sided rule: only Pleasure Ecstasy had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Pleasure Ecstasy clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.26.

Nothing was asked of the other axis, and in fact Embarrassment drifts down from 0.92 to 0.74 (-0.18), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.02 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.03 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.02 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.03, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 30 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.033 before conversion and 0.520 after — it rose by 0.553. Neighbour-to-neighbour the worst pair went 0.045 → 0.520. (The earlier render, with segment 1 left raw, scores 0.393 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pleasure Ecstasy moved +0.263 in the original and +0.662 after conversion — 251 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Embarrassment, -0.182 became -0.253.

Quality. Mean predicted overall quality across the segments went 2.73 → 3.07 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 -0.033 → 0.520 +0.553identity cos neighbours 0.045 → 0.520d_b rescored +0.263 → +0.662d_a rescored -0.182 → -0.253d_a mined -0.177d_b mined 0.257min_cos_consec (site) 0.0159min_cos_anchor (site) -0.0344dataset podcastlang enspeaker 451388total 29.4schain gain +3.6 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, somewhat unclear, normal breath
(embarrassment · measured, normally alert, relaxed, casual) as one of the ba one of the bases for it. And (low mumble) uh Yeah, I just went from there, you
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as embarrassment; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 1.9/10; 6.6s, EN.
451388_00014791 · in -28.4 dBFS · gain +8.3 dB · podcast-02183
(amusement, intoxication altered states of consciousness, elation · brisk, energised, neutral tension, casual) (ahem) talking about. But like you could feel it. Like I've always good been good at feeling things, and like you could just feel what he was, you know what I mean,
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, very dark, very rough, thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as amusement, intoxication altered states of consciousness, elation; style: casual, playful; below-average recording, some background noise; mildly explicit content; genuineness 5.0/6; vocal-burst blend 2.1/10; 9.5s, EN.
451388_00018016 · in -26.0 dBFS · gain +6.0 dB · podcast-04819
(pleasure ecstasy, contempt, intoxication altered states of consciousness · normal-paced, energised, neutral tension, casual) putting out. And it was just like I ball footed smashing a coconut over somebody's head. Like, you know what I mean? Like, if you feel that shit in that moment, do it. And like shit like that is just like I enjoyed. You know what I mean? Like,
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as pleasure ecstasy, contempt, intoxication altered states of consciousness; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 5.3/6; vocal-burst blend 8.1/10; 13.7s, EN.
451388_00019016 · in -27.2 dBFS · gain +7.2 dB · podcast-04800
Interest(unconstrained axis: Thankfulness Gratitude)identity +0.10 emotion 89 %   emotion__B1__T0.25__C0.25__INTERNAL · #14

This chain comes from the one-sided rule: only Interest had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Interest clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.26.

Nothing was asked of the other axis, and in fact Thankfulness Gratitude drifts down from 0.95 to 0.83 (-0.12), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.09 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 38 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.752 before conversion and 0.849 after — it rose by 0.097. Neighbour-to-neighbour the worst pair went 0.758 → 0.764. (The earlier render, with segment 1 left raw, scores 0.685 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.258 in the original and +0.231 after conversion — 89 % of the delta retained, which is most of it. On the other named axis, Thankfulness Gratitude, -0.118 became -0.075.

Quality. Mean predicted overall quality across the segments went 2.55 → 3.01 (+0.46) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.752 → 0.849 +0.097identity cos neighbours 0.758 → 0.764d_b rescored +0.258 → +0.231d_a rescored -0.118 → -0.075d_a mined -0.119d_b mined 0.256min_cos_consec (site) 0.8106min_cos_anchor (site) 0.8085dataset podcastlang enspeaker 275187total 37.3schain gain +3.0 dBseam step 3.6 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child masculine voice · energised, neutral tension, moderately variable, wide pitch range
(thankfulness gratitude, triumph, jealousy and envy · fast, frequent disfluency, slurred, casual) you have a lifetime in front of you. And what I mean by that, what I mean by that, Kevin, is that you have the opportunity right now to do anything you want, to
full caption & clip details
A child masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, dark, slightly rough, thin; slurred, frequent disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, guarded; reads as thankfulness gratitude, triumph, jealousy and envy; style: casual, dramatic; poor recording, quiet background; genuineness 3.9/6; vocal-burst blend 6.1/10; 8.4s, EN.
275187_00208896 · in -29.3 dBFS · gain +9.3 dB · podcast-02729
(hope enthusiasm optimism, impatience and irritability, interest · brisk, some disfluency, average clarity, casual) try anything, to go for anything, to give it your all. What I would tell you to do is figure out what do you think success looks like for you? What
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, impatience and irritability, interest; style: casual, conversational; average recording, some background noise; genuineness 5.2/6; vocal-burst blend 7.0/10; 7.0s, EN.
275187_00209736 · in -28.6 dBFS · gain +8.6 dB · podcast-02732
(interest, hope enthusiasm optimism, elation · brisk, some disfluency, somewhat unclear, authoritative) What and then what are the skills that you need to learn to get there? Who are the people that you need to connect with to get there, and what can you do every single day on repeat, trying and failing, trying and failing, trying and new things, trying and learning, and then continue to do those things to figure out what it is that you want. You want to start a business, try a business tomorrow. You want to build your podcast right now, go find people that you can get on your podcast.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, fairly guarded; reads as interest, hope enthusiasm optimism, elation; style: authoritative, dramatic; below-average recording, quiet background; genuineness 3.1/6; vocal-burst blend 8.0/10; 22.2s, EN.
275187_00210696 · in -27.2 dBFS · gain +7.2 dB · podcast-00479
Pride(unconstrained axis: Disgust)identity +0.02 emotion 239 %   emotion__B1__T0.25__C0.25__INTERNAL · #15

This chain comes from the one-sided rule: only Pride had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Pride clearly present — 0.67, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.32.

Nothing was asked of the other axis, and in fact Disgust drifts down from 0.87 to 0.03 (-0.84), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.14, then +0.18 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 47 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.942 before conversion and 0.960 after — it rose by 0.019. Neighbour-to-neighbour the worst pair went 0.942 → 0.966. (The earlier render, with segment 1 left raw, scores 0.866 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.324 in the original and +0.772 after conversion — 239 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Disgust, -0.845 became +0.138.

Quality. Mean predicted overall quality across the segments went 3.22 → 3.50 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.942 → 0.960 +0.019identity cos neighbours 0.942 → 0.966d_b rescored +0.324 → +0.772d_a rescored -0.845 → +0.138d_a mined -0.845d_b mined 0.324min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00073_S05203total 46.2schain gain +1.0 dBseam step 0.6 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, slightly dark, fairly smooth, balanced body, average recording, quiet background, measured, normally alert
(didactic, monologue) 他们在为这片区域执行任务的见习歧士提供帮助,有专业的人员审查,对当地发生的事件啊,包括他们这个级别呀,还有等等。那么偶尔也会有骑士呢,派驻在教会的内部。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 5.7/10; 15.4s, ZH.
ZH_B00073_S05203_W000007 · in -17.6 dBFS · gain -2.4 dB · emolia-04005
(didactic, monologue) 伦敦街道的教会跟当地最大的保罗大教堂合为一体。当前,卡斯小队全员来到这儿,由神父执行净化仪式,替两位遭到污染的见习骑士驱除污染。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 3.3/10; 14.3s, ZH.
ZH_B00073_S05203_W000008 · in -17.4 dBFS · gain -2.6 dB · emolia-04005
(pride, interest, astonishment surprise · monologue, didactic) 这圣光透过教堂的顶层玻璃垂直降在寒冬,俩人身上一股温暖感在体内散开,寒冬,甚至感觉呀如果自个儿真受伤了,在接受这等圣光的沐浴时,伤势必将在短期内全部恢复。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, interest, astonishment surprise; style: monologue, didactic; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 5.4/10; 16.9s, ZH.
ZH_B00073_S05203_W000009 · in -17.5 dBFS · gain -2.5 dB · emolia-04005
Concentration(unconstrained axis: Intoxication Altered States of Consciousness)identity −0.03 emotion 130 %   emotion__B1__T0.25__C0.25__INTERNAL · #16

This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Concentration around average — 0.46, lower than 54 % of clips in this corpus — and ends with it clearly present at 0.75, higher than 75 % of clips in this corpus. That is a total rise of 0.29.

Nothing was asked of the other axis, and in fact Intoxication Altered States of Consciousness drifts down from 0.97 to 0.15 (-0.83), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.11, then +0.20, then -0.24, then +0.22 — not a clean run: step 3 moves back the other way by 0.24 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.90 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.90 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 35 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.851 before conversion and 0.824 after — it fell by 0.027. Neighbour-to-neighbour the worst pair went 0.852 → 0.824. (The earlier render, with segment 1 left raw, scores 0.790 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.289 in the original and +0.375 after conversion — 130 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Intoxication Altered States of Consciousness, -0.829 became -0.903.

Quality. Mean predicted overall quality across the segments went 3.08 → 3.15 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.851 → 0.824 -0.027identity cos neighbours 0.852 → 0.824d_b rescored +0.289 → +0.375d_a rescored -0.829 → -0.903d_a mined -0.829d_b mined 0.290min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00043_S01426total 33.5schain gain +1.5 dBseam step 0.4 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a child feminine voice · fairly smooth, no background noise, slightly relaxed
(intoxication altered states of consciousness, affection, sexual lust · slow, very low-energy, moderately variable, cartoonish) 在我看来,这些标志要表现的是相对界的消失。
full caption & clip details
A child feminine voice; delivery is very low-energy, slow, slightly relaxed, moderately variable; timbre is cool, dark, fairly smooth, thin; slurred, frequent disfluency, very wide pitch range, audible breath; affect is positive, slightly submissive, neutral openness; reads as intoxication altered states of consciousness, affection, sexual lust; style: cartoonish, storytelling; average recording, no background noise; genuineness 3.2/6; vocal-burst blend 1.4/10; 5.6s, ZH.
ZH_B00043_S01426_W000029 · in -20.2 dBFS · gain +0.2 dB · emolia-03708
(doubt · measured, normally alert, fairly steady, monologue) 从山下开始,攀登的人是怎样聆听基督的教诲的呢?
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as doubt; style: monologue, formal; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.5/10; 4.7s, ZH.
ZH_B00043_S01426_W000030 · in -20.2 dBFS · gain +0.2 dB · emolia-03708
(pain, emotional numbness · fast, normally alert, fairly steady, whispered) 在抵达山顶之前,仰望十字架,十字架的标志和教义看起来就像是最高的终点。
full caption & clip details
A child feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as pain, emotional numbness; style: whispered, monologue; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 2.0/10; 8.4s, ZH.
ZH_B00043_S01426_W000031 · in -20.5 dBFS · gain +0.5 dB · emolia-03708
(measured, normally alert, fairly steady, whispered) 信仰神道的人在攀登途中看见鸟居,认为这里有最高的神。
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is cool, slightly dark, fairly smooth, thin; slurred, no disfluency, narrow pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: whispered, cartoonish; average recording, no background noise; genuineness 1.1/6; vocal-burst blend 0.6/10; 6.5s, ZH.
ZH_B00043_S01426_W000032 · in -19.7 dBFS · gain -0.3 dB · emolia-03708
(measured, normally alert, fairly steady, whispered) 有人从南面向上攀登,途中有佛寺的话,他就会认为寺里有佛,而佛典之中有佛吗?
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: whispered, casual; average recording, no background noise; genuineness 2.2/6; vocal-burst blend 2.4/10; 9.1s, ZH.
ZH_B00043_S01426_W000033 · in -19.4 dBFS · gain -0.6 dB · emolia-03708
Infatuation(unconstrained axis: Shame)identity +0.01 emotion 129 %   emotion__B1__T0.25__C0.25__INTERNAL · #17

This chain comes from the one-sided rule: only Infatuation had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Infatuation around average — 0.48, lower than 52 % of clips in this corpus — and ends with it strongly present at 0.81, higher than 81 % of clips in this corpus. That is a total rise of 0.34.

Nothing was asked of the other axis, and in fact Shame drifts down from 0.87 to 0.31 (-0.56), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.10 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.82 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.82 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 23 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.776 before conversion and 0.787 after — it rose by 0.011. Neighbour-to-neighbour the worst pair went 0.776 → 0.787. (The earlier render, with segment 1 left raw, scores 0.711 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.335 in the original and +0.434 after conversion — 129 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Shame, -0.561 became -0.448.

Quality. Mean predicted overall quality across the segments went 2.97 → 3.08 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.776 → 0.787 +0.011identity cos neighbours 0.776 → 0.787d_b rescored +0.335 → +0.434d_a rescored -0.561 → -0.448d_a mined -0.561d_b mined 0.335min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00044_S07150total 22.6schain gain +1.5 dBseam step 1.3 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, measured, normally alert, slightly relaxed, light breath
(steady, no disfluency, clear, monologue) 加紧了对拉丁美洲的经济侵略和政治渗透。一八二三年美国总统蒙罗发表宣言宣称,美洲是美洲人的美洲,将拉丁美洲视为自己的势力范围。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.7/10; 14.5s, ZH.
ZH_B00044_S07150_W000007 · in -18.6 dBFS · gain -1.4 dB · emolia-03712
(pain · fairly steady, some disfluency, somewhat unclear, monologue) 美国在对拉丁美洲进行经济侵略的同时,还进行武力干涉。
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: monologue, didactic; average recording, no background noise; genuineness 2.1/6; vocal-burst blend 2.2/10; 5.3s, ZH.
ZH_B00044_S07150_W000008 · in -18.9 dBFS · gain -1.1 dB · emolia-03712
(fairly steady, no disfluency, clear, formal) 这就是所谓的金圆、外交和大棒政策。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 2.5/10; 3.2s, ZH.
ZH_B00044_S07150_W000009 · in -20.2 dBFS · gain +0.2 dB · emolia-03712
Amusement(unconstrained axis: Disgust)identity −0.04 emotion 246 %   emotion__B1__T0.25__C0.25__INTERNAL · #18

This chain comes from the one-sided rule: only Amusement had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Amusement clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.25.

Nothing was asked of the other axis, and in fact Disgust barely moves at all, sitting near 1.00 throughout.

It takes 3 clips to get there. Clip to clip the moves are +0.03, then +0.22 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.73 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.66 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.73, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 23 s · de · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.690 before conversion and 0.646 after — it fell by 0.044. Neighbour-to-neighbour the worst pair went 0.690 → 0.646. (The earlier render, with segment 1 left raw, scores 0.502 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Amusement moved +0.259 in the original and +0.637 after conversion — 246 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Disgust, -0.052 became -0.130.

Quality. Mean predicted overall quality across the segments went 2.52 → 2.94 (+0.42) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.690 → 0.646 -0.044identity cos neighbours 0.690 → 0.646d_b rescored +0.259 → +0.637d_a rescored -0.052 → -0.130d_a mined -0.049d_b mined 0.253min_cos_consec (site) 0.6554min_cos_anchor (site) 0.7314dataset podcastlang despeaker 594359total 22.3schain gain +1.8 dBseam step 2.7 dBcrossfades 100/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, quiet background, some disfluency, average clarity, wide pitch range
(disgust, contempt, sourness · brisk, energised, neutral tension, storytelling) habe ich nicht nötig oder so. Ja, genau. Boah, das ist so eklig. Also sorry, aber das ist einfach wirklich uncool. Das ist so richtig schlechter Verlierer. Ich finde, es gibt nichts Abstoßenderes und Unattraktiveres, als ein schlechter Verlierer zu sein, vor allem. Was gab's,
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as disgust, contempt, sourness; style: storytelling, dramatic; good recording, quiet background; mildly explicit content; genuineness 4.3/6; vocal-burst blend 3.0/10; 13.1s, DE.
594359_00229064 · in -26.8 dBFS · gain +6.8 dB · podcast-05794
(confusion, impatience and irritability, sourness · normal-paced, normally alert, slightly relaxed, dramatic) wenn zwischen dir und diesen Typen nichts passiert ist, du Kleine, was ist denn
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as confusion, impatience and irritability, sourness; style: dramatic, conversational; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.0/10; 3.4s, DE.
594359_00230392 · in -22.8 dBFS · gain +2.8 dB · podcast-05811
(amusement, embarrassment, teasing · fast, energised, neutral tension, conversational) (ahem) Maul. Es ist einfach frech und unerzogen und zeigt einfach, was für ein Charakter dieser Mensch hat. Fertig.
full caption & clip details
A young adult feminine voice; delivery is energised, fast, neutral tension, volatile; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as amusement, embarrassment, teasing; style: conversational, casual; below-average recording, quiet background; genuineness 4.9/6; vocal-burst blend 2.5/10; 6.1s, DE.
594359_00231160 · in -19.8 dBFS · gain -0.2 dB · podcast-00402
Relief(unconstrained axis: Pride)identity +0.03 emotion REVERSED   emotion__B1__T0.25__C0.25__INTERNAL · #19

This chain comes from the one-sided rule: only Relief had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Relief around average — 0.47, lower than 53 % of clips in this corpus — and ends with it strongly present at 0.79, higher than 79 % of clips in this corpus. That is a total rise of 0.33.

Nothing was asked of the other axis, and in fact Pride drifts down from 0.96 to 0.69 (-0.27), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.20 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 25 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.741 before conversion and 0.776 after — it rose by 0.035. Neighbour-to-neighbour the worst pair went 0.741 → 0.776. (The earlier render, with segment 1 left raw, scores 0.693 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Relief moved +0.327 in the original and -0.113 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Pride, -0.269 became -0.038.

Quality. Mean predicted overall quality across the segments went 3.00 → 3.10 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.741 → 0.776 +0.035identity cos neighbours 0.741 → 0.776d_b rescored +0.327 → -0.113d_a rescored -0.269 → -0.038d_a mined -0.269d_b mined 0.327min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00017_S07771total 24.6schain gain +3.0 dBseam step 1.2 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed
(pride · measured, steady, almost no disfluency, didactic) 我们概括一下创造的原则有三点。第一点呢就是决定你想要什么。第二点呢要清晰,具体第三点就是把它写下来。
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, almost no disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as pride; style: didactic, monologue; average recording, no background noise; genuineness 2.1/6; vocal-burst blend 2.1/10; 11.6s, ZH.
ZH_B00017_S07771_W000008 · in -17.8 dBFS · gain -2.2 dB · emolia-03442
(normal-paced, fairly steady, little disfluency, formal) 当我们遵循这三点,我们就符合创造原则。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 2.7/6; vocal-burst blend 2.8/10; 3.6s, ZH.
ZH_B00017_S07771_W000009 · in -18.4 dBFS · gain -1.6 dB · emolia-03442
(measured, fairly steady, some disfluency, monologue) 那么我们在生活中就一定能创造出我们想要的结果。如果我们不遵循这三点,我们就很难得到自己想要的结果了。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, authoritative; average recording, no background noise; genuineness 3.1/6; vocal-burst blend 4.4/10; 9.7s, ZH.
ZH_B00017_S07771_W000010 · in -19.1 dBFS · gain -0.9 dB · emolia-03442
Shame(unconstrained axis: Awe)identity +0.24 emotion 106 %   emotion__B1__T0.25__C0.25__INTERNAL · #20

This chain comes from the one-sided rule: only Shame had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Shame clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.27.

Nothing was asked of the other axis, and in fact Awe drifts down from 0.99 to 0.47 (-0.51), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.03, then +0.25, then -0.01 — not a clean run: step 3 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.39 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.18 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.39, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 64 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.451 before conversion and 0.696 after — it rose by 0.245. Neighbour-to-neighbour the worst pair went 0.220 → 0.750. (The earlier render, with segment 1 left raw, scores 0.655 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.269 in the original and +0.285 after conversion — 106 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Awe, -0.513 became -0.481.

Quality. Mean predicted overall quality across the segments went 2.92 → 3.16 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.451 → 0.696 +0.245identity cos neighbours 0.220 → 0.750d_b rescored +0.269 → +0.285d_a rescored -0.513 → -0.481d_a mined -0.512d_b mined 0.269min_cos_consec (site) 0.1765min_cos_anchor (site) 0.3895dataset podcastlang enspeaker 111600total 62.9schain gain +1.4 dBseam step 0.1 dBcrossfades 150/150/100 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-bright, fairly smooth, normally alert, slightly relaxed
(awe, contentment, interest · normal-paced, fairly steady, almost no disfluency, whispered) so true, Rumby. It's similar to history. It's a living, breathing narrative shaped by countless perspectives and interpretations, often immortalized through the strokes of a brush or the chisel of a sculptor.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as awe, contentment, interest; style: whispered, narration; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 2.0/10; 10.8s, EN.
111600_00017240 · in -31.4 dBFS · gain +11.4 dB · podcast-04917
(affection, hope enthusiasm optimism · measured, fairly steady, no disfluency, monologue) We're so excited to introduce our first guest, Parvani and Smiller, who will help us explore these questions about the immortality of art and whether history is an ever-evolving narrative or something more final.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, minimal breath; affect is mildly positive, neutral stance, slightly guarded; reads as affection, hope enthusiasm optimism; style: monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.5/10; 12.6s, EN.
111600_00018344 · in -31.0 dBFS · gain +11.0 dB · podcast-02575
(pride, shame, elation · normal-paced, fairly steady, some disfluency, whispered) As an artist, I guess I'm an aspiring writer and poet, and then professionally I work in arts marketing. I'm a social media manager and like a comms manager consultant. I grew up in Coventry. I was born in South India. Didn't really like I've always loved to write and you know be creative,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; clear, some disfluency, moderate pitch range, light breath; affect is positive, neutral stance, neutral openness; reads as pride, shame, elation; style: whispered, casual; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 4.1/10; 18.4s, EN.
111600_00019920 · in -28.4 dBFS · gain +8.4 dB · podcast-02564
(shame, hope enthusiasm optimism, triumph · brisk, moderately variable, some disfluency) but didn't think that could be a career or something I could realistically do. (low mumble) Um, until I ended up choosing to go study film (ahem) uh production, media production at university. And yeah, here I graduated about two years ago from Coventry, and now I'm just freelancing as a marketing professional and figuring out the whole artist online content creator journey.
full caption & clip details
A young adult somewhat feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, neutral stance, slightly guarded; reads as shame, hope enthusiasm optimism, triumph; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 4.5/10; 21.6s, EN.
111600_00021768 · in -28.9 dBFS · gain +8.9 dB · podcast-02574