k-AB2-k5 — voice-corrected

AB2 at chain length k=5, all corpora, at the mining floor.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_k-AB2-k5.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
80segments re-voiced
0.675 → 0.771median worst-to-anchor identity cosine
105 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Doubt ↓  /  Hope Enthusiasm Optimismidentity +0.01 emotion 90 %   k-AB2-k5 · #1

This chain comes from the two-sided rule: it only counts if both emotions move — Doubt down and Hope Enthusiasm Optimism up — by at least 0.25 each.

The chain starts with Hope Enthusiasm Optimism clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.29.

At the same time Doubt goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.49 (lower than 51 % of clips in this corpus), a change of -0.50. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then -0.17, then +0.12, then +0.12 — not a clean run: step 2 moves back the other way by 0.17 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 70 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.753 before conversion and 0.763 after — it rose by 0.010. Neighbour-to-neighbour the worst pair went 0.748 → 0.746. (The earlier render, with segment 1 left raw, scores 0.499 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.285 in the original and +0.257 after conversion — 90 % of the delta retained, which is essentially all of it. On the other named axis, Doubt, -0.500 became -0.621.

Quality. Mean predicted overall quality across the segments went 2.98 → 3.12 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.753 → 0.763 +0.010identity cos neighbours 0.748 → 0.746d_b rescored +0.285 → +0.257d_a rescored -0.500 → -0.621d_a mined -0.499d_b mined 0.285min_cos_consec (site) 0.8836min_cos_anchor (site) 0.8415dataset emolialang enspeaker EN_B00048_S04457total 68.4schain gain +3.7 dBseam step 2.5 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · slightly cool, fairly smooth, brisk, energised, moderately variable, some disfluency
(doubt, fear, concentration · neutral tension, average clarity, wide pitch range, dramatic) But I worry, just to be clear, we might have now created a problem. It might seem if I play this naively that okay, how do I now actually do math with the number 65 if now Excel displays 65 is an A, let alone B's and C's. So how might a computer
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, slightly guarded; reads as doubt, fear, concentration; style: dramatic, casual; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 2.3/10; 16.4s, EN.
EN_B00048_S04457_W000047 · in -21.4 dBFS · gain +1.4 dB · emolia-01180
(neutral tension, clear, wide pitch range, didactic) Do as you've proposed, have this mapping from numbers to letters, but still support numbers. It feels like we've given something up. Yeah.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: didactic, dramatic; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 0.4/10; 8.4s, EN.
EN_B00048_S04457_W000048 · in -23.8 dBFS · gain +3.8 dB · emolia-01180
(concentration, contentment, interest · slightly relaxed, average clarity, wide pitch range, casual) Okay, so we could perhaps have some kind of prefix like some pattern of zeros and ones. I like this that represents indicates to the computer. Here comes another pattern that represents a letter. Here comes another pattern that represents (ahem) a number or a letter. So not bad. I like that other thoughts. How might a computer distinguish these two? (ahem) Yeah.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration, contentment, interest; style: casual, dramatic; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 3.3/10; 19.6s, EN.
EN_B00048_S04457_W000049 · in -21.1 dBFS · gain +1.1 dB · emolia-01180
(interest, anger, astonishment surprise · fully relaxed, average clarity, very wide pitch range, dramatic) Indeed, and that's spot on. Nothing wrong with what you suggested, but the world generally does just that. The reason we have all of these different file formats in the world like (ahem) jpeg and jiff and pings and word documents dot (ahem) d o s c x and Excel files and so forth is because a bunch of humans got in a room and decided, well, in the context of this type of file,
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, fully relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, very wide pitch range, normal breath; affect is elated, slightly dominant, fairly guarded; reads as interest, anger, astonishment surprise; style: dramatic, casual; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 4.7/10; 18.1s, EN.
EN_B00048_S04457_W000050 · in -21.3 dBFS · gain +1.3 dB · emolia-01180
(hope enthusiasm optimism · neutral tension, very clear, wide pitch range, dramatic) Or really, more specifically, in the context of this type of program, Excel versus Photoshop versus Google Docs or the like.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; very clear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism; style: dramatic, casual; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 2.7/10; 6.7s, EN.
EN_B00048_S04457_W000051 · in -20.9 dBFS · gain +0.9 dB · emolia-01180
Concentration ↓  /  Sournessidentity +0.07 emotion 17 %   k-AB2-k5 · #2

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Sourness up — by at least 0.25 each.

The chain starts with Sourness below average — 0.36, lower than 64 % of clips in this corpus — and ends with it clearly present at 0.66, higher than 66 % of clips in this corpus. That is a total rise of 0.30.

At the same time Concentration goes the other way, from 0.80 (higher than 80 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.08, then -0.04, then +0.07, then +0.18 — not a clean run: step 2 moves back the other way by 0.04 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.80 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.80. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 27 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.793 before conversion and 0.868 after — it rose by 0.075. Neighbour-to-neighbour the worst pair went 0.839 → 0.816. (The earlier render, with segment 1 left raw, scores 0.806 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sourness moved +0.299 in the original and +0.050 after conversion — 17 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Concentration, -0.287 became -0.273.

Quality. Mean predicted overall quality across the segments went 2.95 → 3.01 (+0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.793 → 0.868 +0.075identity cos neighbours 0.839 → 0.816d_b rescored +0.299 → +0.050d_a rescored -0.287 → -0.273d_a mined -0.285d_b mined 0.299min_cos_consec (site) 0.8316min_cos_anchor (site) 0.8034dataset emolialang zhspeaker ZH_B00042_S09822total 25.5schain gain +0.6 dBseam step 1.5 dBcrossfades 100/100/100/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, clear, light breath
(measured, fairly steady, some disfluency, didactic) 对于整个人格的建设具有功能意义。基于性格形成的宽容,总是意味着。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: didactic, monologue; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.7/10; 6.7s, ZH.
ZH_B00042_S09822_W000018 · in -23.3 dBFS · gain +3.3 dB · emolia-03694
(measured, fairly steady, no disfluency, formal) 个体对他人始终抱有积极的尊重,无论对方是谁。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, whispered; average recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.0/10; 5.0s, ZH.
ZH_B00042_S09822_W000019 · in -22.0 dBFS · gain +2.0 dB · emolia-03694
(normal-paced, fairly steady, no disfluency, monologue) 这种尊重可能会延伸到生活方式中的方方面面。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.5/10; 4.5s, ZH.
ZH_B00042_S09822_W000020 · in -22.5 dBFS · gain +2.5 dB · emolia-03694
(measured, steady, no disfluency, whispered) 一些人似乎怀有一种总体上的友善情感,一种真诚的善意。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: whispered, formal; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.0/10; 5.9s, ZH.
ZH_B00042_S09822_W000021 · in -23.1 dBFS · gain +3.0 dB · emolia-03694
(measured, fairly steady, no disfluency, whispered) 外群体的成员对他们来说是新鲜有趣的。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: whispered, monologue; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 3.0/10; 4.0s, ZH.
ZH_B00042_S09822_W000022 · in -23.1 dBFS · gain +3.0 dB · emolia-03694
Contemplation ↓  /  Hope Enthusiasm Optimismidentity −0.01 emotion 122 %   k-AB2-k5 · #3

This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Hope Enthusiasm Optimism up — by at least 0.25 each.

The chain starts with Hope Enthusiasm Optimism clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.32.

At the same time Contemplation goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.58 (higher than 58 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then -0.21, then +0.23, then +0.13 — not a clean run: step 2 moves back the other way by 0.21 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 79 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.841 before conversion and 0.832 after — it fell by 0.009. Neighbour-to-neighbour the worst pair went 0.841 → 0.820. (The earlier render, with segment 1 left raw, scores 0.636 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.317 in the original and +0.386 after conversion — 122 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contemplation, -0.280 became -0.316.

Quality. Mean predicted overall quality across the segments went 3.01 → 3.18 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.841 → 0.832 -0.009identity cos neighbours 0.841 → 0.820d_b rescored +0.317 → +0.386d_a rescored -0.280 → -0.316d_a mined -0.280d_b mined 0.317min_cos_consec (site) 0.8831min_cos_anchor (site) 0.8785dataset emolialang enspeaker EN_B00009_S05665total 78.1schain gain +4.9 dBseam step 1.5 dBcrossfades 150/100/100/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, moderate pitch range
(normal-paced, normally alert, slightly relaxed, casual) You know, when we talk about syndication, basically syndication is a way of raising money from investors to pay for your deal. And (low mumble) uhm, you know, I think that when Brandon and I
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 4.6/6; vocal-burst blend 4.9/10; 12.4s, EN.
EN_B00009_S05665_W000112 · in -24.3 dBFS · gain +4.3 dB · emolia-00436
(pride, infatuation, embarrassment · normal-paced, normally alert, neutral tension, casual) And I actually, I referenced in the book and Brandon referenced it earlier where I was with 30 other multifamily investors. And at that time I was literally the only one in the room that wasn't doing syndications. And I was like, you know, people ask me why, like, why aren't you doing it? And I came up with all kinds of excuses. I'm like, ah, I don't know. Like.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as pride, infatuation, embarrassment; style: casual, conversational; good recording, quiet background; genuineness 4.8/6; vocal-burst blend 10.0/10; 18.4s, EN.
EN_B00009_S05665_W000113 · in -24.2 dBFS · gain +4.2 dB · emolia-00436
(shame, contemplation, disappointment · normal-paced, normally alert, neutral tension, casual) But the truth was I just, I was intimidated by it. I didn't understand it. It seemed complicated. It sounded really complicated when people talked about it. And since then I've learned that it's really not. (low mumble)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as shame, contemplation, disappointment; style: casual, conversational; good recording, quiet background; genuineness 5.0/6; vocal-burst blend 8.9/10; 13.6s, EN.
EN_B00009_S05665_W000114 · in -22.3 dBFS · gain +2.3 dB · emolia-00436
(pride, relief · normal-paced, normally alert, slightly relaxed, casual) There's, there's people, again, it's one of those things, you don't need to do it on your own. There's attorneys that can help you out. There's people that specialize in this. But basically what you're doing is you're forming a general partnership
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as pride, relief; style: casual, whispered; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 4.4/10; 12.8s, EN.
EN_B00009_S05665_W000115 · in -24.4 dBFS · gain +4.4 dB · emolia-00436
(hope enthusiasm optimism, affection, thankfulness gratitude · measured, subdued, slightly relaxed, casual) And it's, it could be one person, it could be multiple people who are raising the capital. They're considered the general partners and you're going out and you're offering equity or participation in the deal to, (low mumble) uh, or ownership to, to, to the people who are putting cash in to help you do the deal. And so that, that allows you as a syndicator.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, minimal breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism, affection, thankfulness gratitude; style: casual, monologue; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 3.8/10; 21.6s, EN.
EN_B00009_S05665_W000116 · in -25.2 dBFS · gain +5.2 dB · emolia-00436
Amusement ↓  /  Hope Enthusiasm Optimismidentity +0.32 emotion 47 %   k-AB2-k5 · #4

This chain comes from the two-sided rule: it only counts if both emotions move — Amusement down and Hope Enthusiasm Optimism up — by at least 0.25 each.

The chain starts with Hope Enthusiasm Optimism clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.29.

At the same time Amusement goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.74 (higher than 74 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.05, then +0.08, then +0.16, then +0.01 — a plateau around step 4, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.02 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.30 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.02, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 76 s · en · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.062 before conversion and 0.382 after — it rose by 0.320. Neighbour-to-neighbour the worst pair went 0.273 → 0.550. (The earlier render, with segment 1 left raw, scores 0.253 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.290 in the original and +0.136 after conversion — 47 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Amusement, -0.259 became -0.030.

Quality. Mean predicted overall quality across the segments went 2.75 → 2.99 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.062 → 0.382 +0.320identity cos neighbours 0.273 → 0.550d_b rescored +0.290 → +0.136d_a rescored -0.259 → -0.030d_a mined -0.255d_b mined 0.290min_cos_consec (site) 0.2975min_cos_anchor (site) 0.0176dataset podcastlang enspeaker 824540total 75.2schain gain +3.2 dBseam step 3.4 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, balanced body
(amusement, teasing, pleasure ecstasy · normal-paced, energised, relaxed, casual) most of the (childlike giggle) time. Uh Leary, who's helping you?
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, relaxed, volatile; timbre is neutral-toned, neutral-bright, very rough, balanced body; slurred, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as amusement, teasing, pleasure ecstasy; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 5.5/6; vocal-burst blend 0.4/10; 5.2s, EN.
824540_00236112 · in -15.4 dBFS · gain -4.6 dB · podcast-00593
(amusement, disgust, embarrassment · brisk, energised, neutral tension, casual) I'm now waiting for Madicus to blow my Discord up about how he runs all his events solo, but more power to him.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as amusement, disgust, embarrassment; style: casual, conversational; good recording, quiet background; genuineness 4.4/6; vocal-burst blend 2.9/10; 6.0s, EN.
824540_00238984 · in -19.5 dBFS · gain -0.5 dB · podcast-00608
(amusement, intoxication altered states of consciousness, teasing · normal-paced, energised, relaxed, casual) (chuckle) So yeah, yeah, our (low mumble) uh our our chat's going on about (ahem) uh staffing for events now. Way to go, Ruth.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, very rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as amusement, intoxication altered states of consciousness, teasing; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 5.7/6; vocal-burst blend 1.5/10; 9.1s, EN.
824540_00239656 · in -17.9 dBFS · gain -2.1 dB · podcast-04589
(hope enthusiasm optimism, elation, contentment · brisk, energised, neutral tension, casual) Triggered. All right, so the biggest question we have we just to make sure everybody is completely clear. How do innkeepers (low mumble) schedule the (ahem) these events, these THQs?
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as hope enthusiasm optimism, elation, contentment; style: casual, conversational; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 5.1/10; 29.6s, EN.
824540_00240716 · in -20.4 dBFS · gain +0.3 dB · podcast-06063
(hope enthusiasm optimism, contentment, elation · brisk, normally alert, slightly relaxed, casual) (ahem) (low mumble) (low mumble) (ahem) Awesome. So (low mumble)
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is positive, neutral stance, neutral openness; reads as hope enthusiasm optimism, contentment, elation; style: casual, conversational; good recording, quiet background; genuineness 1.8/6; vocal-burst blend 5.7/10; 26.0s, EN.
824540_00249482 · in -19.3 dBFS · gain -0.7 dB · podcast-06057
Infatuation ↓  /  Concentrationidentity +0.07 emotion 110 %   k-AB2-k5 · #5

This chain comes from the two-sided rule: it only counts if both emotions move — Infatuation down and Concentration up — by at least 0.25 each.

The chain starts with Concentration below average — 0.38, lower than 62 % of clips in this corpus — and ends with it clearly present at 0.71, higher than 71 % of clips in this corpus. That is a total rise of 0.33.

At the same time Infatuation goes the other way, from 0.70 (higher than 70 % of clips in this corpus) to 0.30 (lower than 70 % of clips in this corpus), a change of -0.39. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.21, then -0.09, then +0.06, then +0.14 — not a clean run: step 2 moves back the other way by 0.09 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.78 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.80 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.78, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 25 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.785 before conversion and 0.853 after — it rose by 0.068. Neighbour-to-neighbour the worst pair went 0.756 → 0.806. (The earlier render, with segment 1 left raw, scores 0.773 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.329 in the original and +0.363 after conversion — 110 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Infatuation, -0.392 became -0.504.

Quality. Mean predicted overall quality across the segments went 2.89 → 3.05 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.785 → 0.853 +0.068identity cos neighbours 0.756 → 0.806d_b rescored +0.329 → +0.363d_a rescored -0.392 → -0.504d_a mined -0.392d_b mined 0.329min_cos_consec (site) 0.8041min_cos_anchor (site) 0.7819dataset emolialang zhspeaker ZH_B00042_S09903total 24.1schain gain +2.3 dBseam step 1.4 dBcrossfades 100/100/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed, no disfluency
(measured, fairly steady, moderate pitch range, authoritative) 凯纳尔甲缔结的灾难性合约的时代。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 3.1/10; 3.3s, ZH.
ZH_B00042_S09903_W000048 · in -20.3 dBFS · gain +0.3 dB · emolia-03695
(measured, fairly steady, moderate pitch range, narration) 贵族以各地区事实上的统治者的姿态出现,并处于争权的有利地位。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 3.1/10; 5.9s, ZH.
ZH_B00042_S09903_W000049 · in -19.5 dBFS · gain -0.5 dB · emolia-03695
(normal-paced, fairly steady, moderate pitch range, formal) 他事实上的新国家起于拿破仑入侵后的局面。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 2.0/10; 3.7s, ZH.
ZH_B00042_S09903_W000050 · in -20.6 dBFS · gain +0.6 dB · emolia-03695
(measured, steady, fairly narrow pitch, formal) 成功的创建一个新的强大的对抗帝国。在合并过程的背景下。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 1.7/10; 6.4s, ZH.
ZH_B00042_S09903_W000051 · in -20.5 dBFS · gain +0.5 dB · emolia-03695
(normal-paced, steady, moderate pitch range, formal) 不列颠在四十多年中束缚了他巩固这样一个新帝国结构的能力。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 1.7/10; 5.5s, ZH.
ZH_B00042_S09903_W000052 · in -20.2 dBFS · gain +0.2 dB · emolia-03695
Interest ↓  /  Embarrassmentidentity +0.60 emotion 92 %   k-AB2-k5 · #6

This chain comes from the two-sided rule: it only counts if both emotions move — Interest down and Embarrassment up — by at least 0.25 each.

The chain starts with Embarrassment clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.35.

At the same time Interest goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.44. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.11, then +0.04, then -0.00 — not a clean run: step 4 moves back the other way by 0.00 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.16 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.11 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.16, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 78 s · en · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.173 before conversion and 0.774 after — it rose by 0.601. Neighbour-to-neighbour the worst pair went 0.280 → 0.726. (The earlier render, with segment 1 left raw, scores 0.597 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.347 in the original and +0.318 after conversion — 92 % of the delta retained, which is essentially all of it. On the other named axis, Interest, -0.452 became -0.499.

Quality. Mean predicted overall quality across the segments went 2.78 → 3.03 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.173 → 0.774 +0.601identity cos neighbours 0.280 → 0.726d_b rescored +0.347 → +0.318d_a rescored -0.452 → -0.499d_a mined -0.443d_b mined 0.347min_cos_consec (site) 0.1101min_cos_anchor (site) 0.1644dataset podcastlang enspeaker 811472total 76.9schain gain +2.7 dBseam step 2.9 dBcrossfades 150/100/100/100 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, some disfluency, average clarity
(interest, concentration, contemplation · normal-paced, subdued, slightly relaxed, whispered) But I do think even though actual physical war is technically off of the page, we see a lot of interpersonal war in these five chapters. So we have like the the silent war between Darcy and Wicca. We have the war between like the Bingley sisters and the Bennett sisters, and there's all this like interpersonal conflict that's happening that I feel like we can relate to the theme really well.
full caption & clip details
A young adult feminine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, minimal breath; affect is mildly positive, neutral stance, neutral openness; reads as interest, concentration, contemplation; style: whispered, casual; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 2.0/10; 23.4s, EN.
811472_00187112 · in -38.7 dBFS · gain +18.7 dB · podcast-04553
(amusement, infatuation, doubt · normal-paced, normally alert, slightly relaxed, casual) I was thinking about it specifically in terms of war because of today anyway, but I probably still would have termed Elizabeth and Darcy's conversation as verbal sparring is what I wrote down. And we're thinking about like both how war shows up in our language, but also how war showed up in the theme. I was thinking about that where they're like, again, throwing jabs back and forth, (ahem) um, at least from Elizabeth's perspective. Darcy's just trying to engage in the conversation. He's not trying to like win anything. He's just trying to figure out what's going on. But (breathy giggle) Elizabeth thinks
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as amusement, infatuation, doubt; style: casual, conversational; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 3.4/10; 28.0s, EN.
811472_00189448 · in -34.0 dBFS · gain +14.1 dB · podcast-04558
(amusement, sexual lust, pleasure ecstasy · normal-paced, normally alert, relaxed, casual) throwing like these punches left and right, like boom, uppercut, boom, point Elizabeth. It's just
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as amusement, sexual lust, pleasure ecstasy; style: casual, conversational; average recording, no background noise; mildly explicit content; genuineness 4.5/6; vocal-burst blend 4.9/10; 5.8s, EN.
811472_00192352 · in -29.1 dBFS · gain +9.2 dB · podcast-02161
(amusement, embarrassment, sexual lust · brisk, energised, neutral tension, casual) (breathy giggle) not See, yet again, this is why Darcy is an honorary queer, 'cause he just doesn't know what g what's going on. He likes this girl and he's like, Yeah, okay. I'm
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, volatile; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, very vulnerable; reads as amusement, embarrassment, sexual lust; style: casual, conversational; below-average recording, quiet background; mildly explicit content; genuineness 4.5/6; vocal-burst blend 2.4/10; 9.5s, EN.
811472_00192928 · in -33.7 dBFS · gain +13.7 dB · podcast-02175
(embarrassment, confusion, doubt · normal-paced, normally alert, relaxed, casual) trying to follow your conversation. I don't really understand what's happening (low mumble) here.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as embarrassment, confusion, doubt; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 3.1/6; vocal-burst blend 0.8/10; 10.8s, EN.
811472_00193872 · in -34.2 dBFS · gain +14.2 dB · podcast-04558
Concentration ↓  /  Emotional Numbnessidentity +0.06 emotion 80 %   k-AB2-k5 · #7

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness around average — 0.57, higher than 57 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.35.

At the same time Concentration goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are -0.01, then +0.19, then -0.06, then +0.23 — not a clean run: step 1 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 46 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.785 before conversion and 0.842 after — it rose by 0.057. Neighbour-to-neighbour the worst pair went 0.855 → 0.889. (The earlier render, with segment 1 left raw, scores 0.737 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.346 in the original and +0.276 after conversion — 80 % of the delta retained, which is most of it. On the other named axis, Concentration, -0.271 became -0.212.

Quality. Mean predicted overall quality across the segments went 2.89 → 3.11 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.785 → 0.842 +0.057identity cos neighbours 0.855 → 0.889d_b rescored +0.346 → +0.276d_a rescored -0.271 → -0.212d_a mined -0.271d_b mined 0.346min_cos_consec (site) 0.9112min_cos_anchor (site) 0.9196dataset emolialang enspeaker EN_w6hfhJqGSKAtotal 44.2schain gain +1.5 dBseam step 0.8 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert, slightly relaxed
(concentration, interest · fairly steady, almost no disfluency, minimal breath, newsreading) The traction alternator usually incorporates integral silicon diode rectifiers to provide the traction motors with up to 1200 V DC DC traction, which is used directly, or the common inverter bus' AC traction, which is first inverted from DC to three-phase AC.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, interest; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.1/10; 15.6s, EN.
EN_w6hfhJqGSKA_W000101 · in -18.2 dBFS · gain -1.8 dB · emolia-00592
(interest, elation · steady, almost no disfluency, light breath, newsreading) The first diesel-electric locomotives, and many of those still in service, used DC generators as, before silicon power electronics, it was easier to control the speed of DC traction motors
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, elation; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.1/10; 10.2s, EN.
EN_w6hfhJqGSKA_W000102 · in -17.8 dBFS · gain -2.2 dB · emolia-00592
(fairly steady, no disfluency, light breath, formal) Most of these had two generators, one to generate the excitation current for a larger main generator
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.5/10; 5.1s, EN.
EN_w6hfhJqGSKA_W000103 · in -16.2 dBFS · gain -3.8 dB · emolia-00592
(fairly steady, no disfluency, light breath, formal) Optionally, the generator also supplies head-end power, HEP, or power for electric train heating
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 1.0/10; 5.1s, EN.
EN_w6hfhJqGSKA_W000104 · in -16.7 dBFS · gain -3.3 dB · emolia-00592
(emotional numbness · fairly steady, no disfluency, light breath, newsreading) The HEP option requires a constant engine speed, typically 900 RPM for a 480 V 60 Hz HEP application, even when the locomotive is not moving
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.3/10; 9.0s, EN.
EN_w6hfhJqGSKA_W000105 · in -17.3 dBFS · gain -2.7 dB · emolia-00592
Sexual Lust ↓  /  Elationidentity +0.14 emotion 110 %   k-AB2-k5 · #8

This chain comes from the two-sided rule: it only counts if both emotions move — Sexual Lust down and Elation up — by at least 0.25 each.

The chain starts with Elation clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.26.

At the same time Sexual Lust goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.37. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.16, then +0.12, then -0.08, then +0.06 — not a clean run: step 3 moves back the other way by 0.08 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.68 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.71 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.68, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 48 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.628 before conversion and 0.766 after — it rose by 0.138. Neighbour-to-neighbour the worst pair went 0.743 → 0.728. (The earlier render, with segment 1 left raw, scores 0.695 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Elation moved +0.261 in the original and +0.286 after conversion — 110 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Sexual Lust, -0.368 became -0.143.

Quality. Mean predicted overall quality across the segments went 2.80 → 3.06 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.628 → 0.766 +0.138identity cos neighbours 0.743 → 0.728d_b rescored +0.261 → +0.286d_a rescored -0.368 → -0.143d_a mined -0.368d_b mined 0.261min_cos_consec (site) 0.7141min_cos_anchor (site) 0.6788dataset emolialang enspeaker EN_B00063_S00114total 46.6schain gain +0.3 dBseam step 2.2 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned
(sexual lust, pain · normal-paced, normally alert, neutral tension, casual) (ahem) In Chinese medicine, in Qigong, in martial arts, so many people.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as sexual lust, pain; style: casual, conversational; average recording, no background noise; mildly explicit content; genuineness 4.2/6; vocal-burst blend 2.0/10; 5.6s, EN.
EN_B00063_S00114_W000031 · in -18.8 dBFS · gain -1.2 dB · emolia-01451
(triumph, pain, pride · normal-paced, normally alert, slightly relaxed, casual) Bring a punch to us and instead of like resist it, we actually yield to it.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as triumph, pain, pride; style: casual, monologue; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 1.4/10; 5.1s, EN.
EN_B00063_S00114_W000032 · in -18.1 dBFS · gain -1.9 dB · emolia-01451
(contentment, pleasure ecstasy, thankfulness gratitude · slow, lethargic, relaxed, casual) Nice. Let's open the hands and open the eyes. Beautiful. Thank you guys so much. Welcome and thank you. And I'll see you next time or in class. Have a great, (ahem) uh, great time and great holiday for all of us. Bye now.
full caption & clip details
An elderly masculine voice; delivery is lethargic, slow, relaxed, moderately variable; timbre is neutral-toned, slightly bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as contentment, pleasure ecstasy, thankfulness gratitude; style: casual, conversational; below-average recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.1/10; 17.2s, EN.
EN_B00063_S00114_W000033 · in -19.5 dBFS · gain -0.5 dB · emolia-01451
(affection, thankfulness gratitude, elation · normal-paced, normally alert, slightly relaxed, casual) Uh, (low mumble) there's winter and it's cold. (low mumble) Uh, we say to connect with the energies of gratitude, of love, of, of joy. And if you, (ahem) and we do it also already in our culture, right?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as affection, thankfulness gratitude, elation; style: casual, conversational; good recording, quiet background; genuineness 2.2/6; vocal-burst blend 3.1/10; 11.1s, EN.
EN_B00063_S00114_W000034 · in -17.1 dBFS · gain -2.9 dB · emolia-01451
(elation, hope enthusiasm optimism, pleasure ecstasy · normal-paced, normally alert, slightly relaxed, casual) Yeah, people that likes to go bungee jumping and do, you know, they're very fiery. They're looking for an exhilaration.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as elation, hope enthusiasm optimism, pleasure ecstasy; style: casual, conversational; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 3.1/10; 8.5s, EN.
EN_B00063_S00114_W000035 · in -17.1 dBFS · gain -2.9 dB · emolia-01451
Hope Enthusiasm Optimism ↓  /  Longingidentity −0.07 emotion 75 %   k-AB2-k5 · #9

This chain comes from the two-sided rule: it only counts if both emotions move — Hope Enthusiasm Optimism down and Longing up — by at least 0.25 each.

The chain starts with Longing around average — 0.43, lower than 57 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.44.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are -0.15, then +0.23, then +0.16, then +0.20 — not a clean run: step 1 moves back the other way by 0.15 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 41 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.795 before conversion and 0.724 after — it fell by 0.071. Neighbour-to-neighbour the worst pair went 0.795 → 0.732. (The earlier render, with segment 1 left raw, scores 0.742 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.442 in the original and +0.333 after conversion — 75 % of the delta retained, which is most of it. On the other named axis, Hope Enthusiasm Optimism, -0.298 became -0.234.

Quality. Mean predicted overall quality across the segments went 2.66 → 2.91 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.795 → 0.724 -0.071identity cos neighbours 0.795 → 0.732d_b rescored +0.442 → +0.333d_a rescored -0.298 → -0.234d_a mined -0.299d_b mined 0.442min_cos_consec (site) 0.8424min_cos_anchor (site) 0.8514dataset emolialang enspeaker EN_B00042_S02580total 39.5schain gain +1.4 dBseam step 1.6 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a middle-aged feminine voice · moderately variable, light breath
(hope enthusiasm optimism, contentment, relief · normal-paced, normally alert, slightly relaxed, conversational) Thank you, Mr. Mondale. Our time is up for this round. We go into the second round of our questioning. Begin again (low mumble) with Jim Weehart. Jim?
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, neutral stance, neutral openness; reads as hope enthusiasm optimism, contentment, relief; style: conversational, playful; below-average recording, quiet background; genuineness 3.5/6; vocal-burst blend 2.1/10; 7.0s, EN.
EN_B00042_S02580_W000006 · in -19.9 dBFS · gain -0.1 dB · emolia-01049
(thankfulness gratitude, affection, embarrassment · brisk, normally alert, neutral tension, casual) I'm sorry to do this, but I really must talk to the audience. You're all invited guests. I know I'm wasting time in talking to you, but it really is very unfair of you to applaud sometimes louder, (low mumble) uh, less loud, and I ask you as people who are invited here and polite people to refrain. We have our time now for rebuttal. President.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as thankfulness gratitude, affection, embarrassment; style: casual, conversational; average recording, some background noise; genuineness 4.1/6; vocal-burst blend 8.0/10; 18.4s, EN.
EN_B00042_S02580_W000007 · in -19.5 dBFS · gain -0.5 dB · emolia-01049
(amusement · brisk, normally alert, neutral tension, casual) It's something we all want to do. Two and three questions is part one and two and three is part two. Having said that, Fred, it's yours.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as amusement; style: casual, conversational; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 4.1/10; 5.4s, EN.
EN_B00042_S02580_W000008 · in -21.4 dBFS · gain +1.4 dB · emolia-01049
(helplessness, sadness · brisk, energised, slightly relaxed, casual) And now it is time for our rebuttal for this period, Mr. President.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, thin; average clarity, no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as helplessness, sadness; style: casual, storytelling; below-average recording, quiet background; genuineness 3.8/6; vocal-burst blend 4.0/10; 3.1s, EN.
EN_B00042_S02580_W000009 · in -21.6 dBFS · gain +1.6 dB · emolia-01049
(brisk, normally alert, neutral tension, casual) We now start our final round (ahem) of questions. We do want to have time for your rebuttal. (ahem) We start with Diane. Diane Sula.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 5.1/10; 6.2s, EN.
EN_B00042_S02580_W000010 · in -20.6 dBFS · gain +0.6 dB · emolia-01049
Fear ↓  /  Emotional Numbnessidentity −0.03 emotion 121 %   k-AB2-k5 · #10

This chain comes from the two-sided rule: it only counts if both emotions move — Fear down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.36.

At the same time Fear goes the other way, from 0.79 (higher than 79 % of clips in this corpus) to 0.30 (lower than 70 % of clips in this corpus), a change of -0.49. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.04, then +0.05, then +0.06, then +0.21 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 58 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.857 before conversion and 0.823 after — it fell by 0.034. Neighbour-to-neighbour the worst pair went 0.796 → 0.799. (The earlier render, with segment 1 left raw, scores 0.619 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.358 in the original and +0.432 after conversion — 121 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Fear, -0.493 became -0.388.

Quality. Mean predicted overall quality across the segments went 3.05 → 3.20 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.857 → 0.823 -0.034identity cos neighbours 0.796 → 0.799d_b rescored +0.358 → +0.432d_a rescored -0.493 → -0.388d_a mined -0.493d_b mined 0.358min_cos_consec (site) 0.8896min_cos_anchor (site) 0.9100dataset emolialang enspeaker EN_6PPnj7y5okutotal 56.6schain gain +2.1 dBseam step 0.2 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fairly steady, no disfluency, formal, authoritative) The chefs who had accompanied Nawab Wajid Ali Shah tried various combinations and experiments to enhance the taste of biryani
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 7.1s, EN.
EN_6PPnj7y5oku_W000093 · in -14.2 dBFS · gain -5.8 dB · emolia-02015
(concentration, interest · steady, almost no disfluency, newsreading, authoritative) The Calcutta biryani is much lighter on spices. The marinade primarily uses nutmeg, cinnamon, mace along with cloves and cardamom in the dahi-based marinade for the meat which is cooked separately from rice. This combination of spices gives it a distinct flavor as compared to other styles of biryani.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, interest; style: newsreading, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 17.5s, EN.
EN_6PPnj7y5oku_W000095 · in -15.0 dBFS · gain -5.0 dB · emolia-02015
(fairly steady, no disfluency, formal, narration) The rice is flavored with ketaki water or rose water along with saffron to give it flavor and light yellowish color.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.3/10; 8.1s, EN.
EN_6PPnj7y5oku_W000096 · in -14.2 dBFS · gain -5.8 dB · emolia-02015
(fairly steady, no disfluency, formal, newsreading) Amber – Vaniyambadi biryani is a type of biryani cooked in neighbouring towns of Amber and Vaniyambadi in the Vellore district in the northeastern part of Tamil Nadu, which has a high Muslim population.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.9s, EN.
EN_6PPnj7y5oku_W000097 · in -14.8 dBFS · gain -5.2 dB · emolia-02015
(emotional numbness, disgust, longing · fairly steady, no disfluency, newsreading, formal) It was introduced by the Nawabs of Arkot who once ruled the place.The amber, vaniambadi biryani is accompanied with dhulcha, a sour brinjal curry and pachadi or ritha, which is sliced onions mixed with plain curd, tomato, chillies and salt
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, disgust, longing; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 13.8s, EN.
EN_6PPnj7y5oku_W000098 · in -15.2 dBFS · gain -4.8 dB · emolia-02015
Anger ↓  /  Fearidentity +0.31 emotion 112 %   k-AB2-k5 · #11

This chain comes from the two-sided rule: it only counts if both emotions move — Anger down and Fear up — by at least 0.25 each.

The chain starts with Fear below average — 0.35, lower than 65 % of clips in this corpus — and ends with it strongly present at 0.79, higher than 79 % of clips in this corpus. That is a total rise of 0.44.

At the same time Anger goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.08, then +0.09, then +0.03 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.42 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.40 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.42, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 27 s · sv · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.512 before conversion and 0.821 after — it rose by 0.309. Neighbour-to-neighbour the worst pair went 0.429 → 0.743. (The earlier render, with segment 1 left raw, scores 0.707 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.467 in the original and +0.525 after conversion — 112 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Anger, -0.258 became -0.076.

Quality. Mean predicted overall quality across the segments went 2.84 → 2.97 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.512 → 0.821 +0.309identity cos neighbours 0.429 → 0.743d_b rescored +0.467 → +0.525d_a rescored -0.258 → -0.076d_a mined -0.251d_b mined 0.441min_cos_consec (site) 0.4017min_cos_anchor (site) 0.4172dataset podcastlang svspeaker 627930total 25.4schain gain +2.3 dBseam step 2.2 dBcrossfades 100/150/100/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, average clarity
(slightly relaxed, fairly steady, some disfluency, casual) är ju avspöket. Jag tycker ändå vi tar vägen. Även om det är lite risky.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; good recording, quiet background; genuineness 4.0/6; vocal-burst blend 3.1/10; 4.2s, SV.
627930_00220296 · in -24.2 dBFS · gain +4.2 dB · podcast-03119
(slightly relaxed, fairly steady, no disfluency, casual) Om folk ska ögonbinden blir nog väldigt svårt att gå i dungel.
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, storytelling; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 3.8/10; 3.3s, SV.
627930_00220816 · in -18.3 dBFS · gain -1.7 dB · podcast-03124
(intoxication altered states of consciousness, relief, amusement · neutral tension, moderately variable, some disfluency, casual) Axel, du vet inte hur jobbigt är att gå i vegetationen för du är storstbor. (low mumble) Det kommer ju ta till fem gånger så långt tid.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as intoxication altered states of consciousness, relief, amusement; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 6.0/6; vocal-burst blend 6.1/10; 6.7s, SV.
627930_00222192 · in -21.6 dBFS · gain +1.6 dB · podcast-03122
(embarrassment, amusement · neutral tension, moderately variable, some disfluency, casual) Jag har varit med i mulled och strövarna. (low mumble) Det är ju ta till fem gånger så lång tid att traska igenom en dungel var så traska på vägen.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as embarrassment, amusement; style: casual, conversational; average recording, quiet background; genuineness 5.5/6; vocal-burst blend 2.2/10; 8.3s, SV.
627930_00223287 · in -20.4 dBFS · gain +0.4 dB · podcast-03118
(slightly relaxed, fairly steady, some disfluency, casual) pistolen i en handen. (low mumble) Och går rakt genom skuggan.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 2.4/10; 3.5s, SV.
627930_00225735 · in -19.8 dBFS · gain -0.2 dB · podcast-03142
Longing ↓  /  Emotional Numbnessidentity +0.07 emotion 123 %   k-AB2-k5 · #12

This chain comes from the two-sided rule: it only counts if both emotions move — Longing down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.29.

At the same time Longing goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.75 (higher than 75 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then -0.16, then +0.13, then +0.16 — not a clean run: step 2 moves back the other way by 0.16 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.71 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.79 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.71, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 49 s · ko · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.696 before conversion and 0.769 after — it rose by 0.072. Neighbour-to-neighbour the worst pair went 0.674 → 0.748. (The earlier render, with segment 1 left raw, scores 0.661 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.291 in the original and +0.358 after conversion — 123 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Longing, -0.253 became -0.054.

Quality. Mean predicted overall quality across the segments went 2.95 → 3.18 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.696 → 0.769 +0.072identity cos neighbours 0.674 → 0.748d_b rescored +0.291 → +0.358d_a rescored -0.253 → -0.054d_a mined -0.253d_b mined 0.291min_cos_consec (site) 0.7898min_cos_anchor (site) 0.7148dataset emolialang kospeaker KO_fe3XTX772jAtotal 47.9schain gain +2.1 dBseam step 3.0 dBcrossfades 150/150/100/100 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: an elderly masculine voice · very low-energy, frequent disfluency
(longing, distress, helplessness · measured, relaxed, moderately variable, casual) 그래가지고, 뭐, 게스트 통도 충분한가, 뭐, 다 체크해야 되고. (low mumble) 응. 저거 좀 해보면, 아이고, 아이고, 더 좋은데.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, slightly thin; slurred, frequent disfluency, wide pitch range, audible breath; affect is positive, neutral stance, neutral openness; reads as longing, distress, helplessness; style: casual, playful; below-average recording, quiet background; genuineness 4.5/6; vocal-burst blend 1.0/10; 11.5s, KO.
KO_fe3XTX772jA_W000171 · in -18.8 dBFS · gain -1.2 dB · emolia-03240
(helplessness, longing, pain · measured, relaxed, fairly steady, whispered) 재밌는 추억을 만들려고 노력을 많이 했는데, 제가 이제 뉴스 하트를 알고부터는 너무 바빠져가지고.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is warm, slightly dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, neutral stance, fairly guarded; reads as helplessness, longing, pain; style: whispered, monologue; below-average recording, quiet background; genuineness 2.7/6; vocal-burst blend 3.1/10; 9.4s, KO.
KO_fe3XTX772jA_W000172 · in -22.4 dBFS · gain +2.4 dB · emolia-03240
(sexual lust, longing, pain · slow, relaxed, fairly steady, casual) 고구마도 주고, 뭐. 이게 딱 가능하면 건강식을 해야 되니까. (ahem)
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, very dark, fairly smooth, slightly thin; slurred, frequent disfluency, narrow pitch range, audible breath; affect is mildly negative, submissive, neutral openness; reads as sexual lust, longing, pain; style: casual, conversational; below-average recording, quiet background; genuineness 4.8/6; vocal-burst blend 1.2/10; 6.8s, KO.
KO_fe3XTX772jA_W000173 · in -22.4 dBFS · gain +2.4 dB · emolia-03240
(sexual lust, anger, contemplation · slow, relaxed, fairly steady, whispered) 그래가지고, 이제, 오늘 밤에 또, 강의 끝나면 또, 응, (ahem) 개 밥 준비해야 돼. 네. 그리고 이제, 다, 어떻게, 반드시 운동 시켜야 돼. 응, 시키고. 그런데,
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is warm, dark, rough, slightly thin; slurred, frequent disfluency, narrow pitch range, audible breath; affect is mildly negative, submissive, fairly guarded; reads as sexual lust, anger, contemplation; style: whispered, monologue; below-average recording, quiet background; genuineness 3.4/6; vocal-burst blend 3.7/10; 16.2s, KO.
KO_fe3XTX772jA_W000174 · in -24.7 dBFS · gain +4.7 dB · emolia-03240
(slow, slightly relaxed, steady, didactic) 그런 것이 귀찮다는 것은 사실이에요.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, rough, balanced body; crisply articulate, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, formal; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.4/10; 4.7s, KO.
KO_fe3XTX772jA_W000175 · in -19.5 dBFS · gain -0.5 dB · emolia-03240
Contempt ↓  /  Shameidentity +0.47 emotion 6 %   k-AB2-k5 · #13

This chain comes from the two-sided rule: it only counts if both emotions move — Contempt down and Shame up — by at least 0.25 each.

The chain starts with Shame around average — 0.58, higher than 58 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.29.

At the same time Contempt goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.35. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.05, then +0.00, then +0.11, then +0.13 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.34 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.33 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.34, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 82 s · nl · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.310 before conversion and 0.785 after — it rose by 0.475. Neighbour-to-neighbour the worst pair went 0.289 → 0.696. (The earlier render, with segment 1 left raw, scores 0.555 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.287 in the original and +0.016 after conversion — 6 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Contempt, -0.356 became -0.120.

Quality. Mean predicted overall quality across the segments went 2.56 → 3.34 (+0.78) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.310 → 0.785 +0.475identity cos neighbours 0.289 → 0.696d_b rescored +0.287 → +0.016d_a rescored -0.356 → -0.120d_a mined -0.355d_b mined 0.287min_cos_consec (site) 0.3273min_cos_anchor (site) 0.3382dataset podcastlang nlspeaker 66296total 80.9schain gain +2.6 dBseam step 2.0 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · slightly rough
(contempt, jealousy and envy · normal-paced, very low-energy, neutral tension, casual) Wat leer (ahem) ik er van Jeffra (low mumble) Stemband?
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, neutral tension, fairly steady; timbre is slightly cool, dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, jealousy and envy; style: casual, conversational; below-average recording, quiet background; genuineness 5.5/6; vocal-burst blend 9.0/10; 23.4s, NL.
66296_00063000 · in -38.9 dBFS · gain +18.9 dB · podcast-02794
(jealousy and envy, anger, bitterness · measured, subdued, slightly relaxed, monologue) Alles kan hier gezegd worden. En soms zijn er ook dingen.
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, anger, bitterness; style: monologue, cartoonish; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 5.6/10; 24.4s, NL.
66296_00074024 · in -41.4 dBFS · gain +21.4 dB · podcast-02786
(disgust, contemplation · normal-paced, normally alert, slightly relaxed, monologue) Die blijven ook hier. Maar soms zijn er dingen waarvan ik denk het is goed als we dat ook elders nog eens neerleggen. Maar dat doe ik nooit om dat eerst met jou kort te sluiten. Dat is ook zo'n kader weer.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, thin; slurred, some disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, contemplation; style: monologue, authoritative; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 2.1/10; 9.5s, NL.
66296_00076464 · in -42.1 dBFS · gain +22.1 dB · podcast-05247
(thankfulness gratitude, jealousy and envy, fatigue exhaustion · fast, normally alert, neutral tension, casual) Dat geeft ook veiligheid. Ik kan me ook herinneren, ik ben natuurlijk ook een aantal van jouw trainingen aanwezig geweest. Dat je die ook begint, letterlijk met het uitvragen van dat kader.
full caption & clip details
A young adult somewhat masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is slightly cool, very dark, slightly rough, thin; slurred, some disfluency, moderate pitch range, audible breath; affect is mildly positive, neutral stance, slightly guarded; reads as thankfulness gratitude, jealousy and envy, fatigue exhaustion; style: casual, dramatic; poor recording, some background noise; genuineness 3.8/6; vocal-burst blend 6.5/10; 8.2s, NL.
66296_00077416 · in -44.8 dBFS · gain +24.8 dB · podcast-01101
(normal-paced, very low-energy, neutral tension, casual) veiligheid. Want ik kan zeggen, van dit vind ik belangrijk in onze samenwerking. Maar ik haal het uit de groep. Wat is voor jullie helpend, op de manier zullen zij met elkaar gaan samenwerken om het komende dagdeel met elkaar ergens te komen. En er komen altijd hele mooie dingen uit. En dan is het ook iets van een groep. Dus
full caption & clip details
An elderly masculine voice; delivery is very low-energy, normal-paced, neutral tension, fairly steady; timbre is slightly cool, dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; below-average recording, quiet background; genuineness 4.4/6; vocal-burst blend 6.5/10; 16.1s, NL.
66296_00079440 · in -42.0 dBFS · gain +22.0 dB · podcast-01073
Relief ↓  /  Astonishment Surpriseidentity +0.16 emotion 121 %   k-AB2-k5 · #14

This chain comes from the two-sided rule: it only counts if both emotions move — Relief down and Astonishment Surprise up — by at least 0.25 each.

The chain starts with Astonishment Surprise below average — 0.36, lower than 64 % of clips in this corpus — and ends with it clearly present at 0.66, higher than 66 % of clips in this corpus. That is a total rise of 0.30.

At the same time Relief goes the other way, from 0.85 (higher than 85 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.16, then +0.13, then +0.23, then -0.23 — not a clean run: step 4 moves back the other way by 0.23 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.67 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.77 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.67, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 29 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.654 before conversion and 0.810 after — it rose by 0.157. Neighbour-to-neighbour the worst pair went 0.701 → 0.777. (The earlier render, with segment 1 left raw, scores 0.774 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.300 in the original and +0.364 after conversion — 121 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Relief, -0.351 became -0.335.

Quality. Mean predicted overall quality across the segments went 3.08 → 3.21 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.654 → 0.810 +0.157identity cos neighbours 0.701 → 0.777d_b rescored +0.300 → +0.364d_a rescored -0.351 → -0.335d_a mined -0.333d_b mined 0.300min_cos_consec (site) 0.7739min_cos_anchor (site) 0.6726dataset emolialang zhspeaker ZH_B00039_S09706total 27.8schain gain +1.8 dBseam step 1.1 dBcrossfades 150/150/100/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, slightly relaxed, fairly steady, moderate pitch range
(measured, normally alert, no disfluency, formal) 当天晚上,女仆伺候他他发现这姑娘在哭。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 5.0/10; 3.7s, ZH.
ZH_B00039_S09706_W000043 · in -22.5 dBFS · gain +2.5 dB · emolia-03664
(contemplation · measured, normally alert, little disfluency, narration) 他这时厌恶艾丽莎,刚刚还粗暴的对待过他,可是又请求他原谅艾丽莎哭得更凶了。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as contemplation; style: narration, monologue; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 7.0/10; 8.0s, ZH.
ZH_B00039_S09706_W000044 · in -21.3 dBFS · gain +1.3 dB · emolia-03664
(measured, normally alert, little disfluency, narration) 他说,如果女主人允许,他将把他的不幸全都清吐出来。说吧。德莱娜夫人答道。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; no dominant emotion; style: narration, storytelling; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 5.3/10; 8.0s, ZH.
ZH_B00039_S09706_W000045 · in -20.6 dBFS · gain +0.6 dB · emolia-03664
(pain, teasing · measured, very low-energy, no disfluency, storytelling) 谁拒绝您德莱纳夫人喘不过气来了。
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as pain, teasing; style: storytelling, narration; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 6.4/10; 3.4s, ZH.
ZH_B00039_S09706_W000046 · in -19.6 dBFS · gain -0.4 dB · emolia-03664
(fast, normally alert, no disfluency, storytelling) 德雷娜夫人不再听女仆说了,她大喜过望,几乎丧失了理智。
full caption & clip details
An adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: storytelling, narration; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 4.5/10; 5.5s, ZH.
ZH_B00039_S09706_W000047 · in -17.1 dBFS · gain -2.9 dB · emolia-03664
Relief ↓  /  Disgustidentity +0.06 emotion 35 %   k-AB2-k5 · #15

This chain comes from the two-sided rule: it only counts if both emotions move — Relief down and Disgust up — by at least 0.25 each.

The chain starts with Disgust around average — 0.53, higher than 53 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.37.

At the same time Relief goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.38 (lower than 62 % of clips in this corpus), a change of -0.53. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.12, then -0.12, then +0.21, then +0.16 — not a clean run: step 2 moves back the other way by 0.12 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 38 s · en · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.718 before conversion and 0.778 after — it rose by 0.060. Neighbour-to-neighbour the worst pair went 0.788 → 0.804. (The earlier render, with segment 1 left raw, scores 0.668 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disgust moved +0.366 in the original and +0.127 after conversion — 35 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Relief, -0.548 became +0.264.

Quality. Mean predicted overall quality across the segments went 2.90 → 3.02 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.718 → 0.778 +0.060identity cos neighbours 0.788 → 0.804d_b rescored +0.366 → +0.127d_a rescored -0.548 → +0.264d_a mined -0.527d_b mined 0.367min_cos_consec (site) 0.8883min_cos_anchor (site) 0.8342dataset podcastlang enspeaker 21786total 36.6schain gain +2.3 dBseam step 1.2 dBcrossfades 100/100/150/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(relief · normal-paced, fairly steady, some disfluency, casual) So the br the British government has never apologised for this.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as relief; style: casual, conversational; good recording, no background noise; genuineness 3.9/6; vocal-burst blend 0.8/10; 3.3s, EN.
21786_00258183 · in -25.1 dBFS · gain +5.1 dB · podcast-04630
(measured, fairly steady, frequent disfluency, casual) you mentioned briefly, (low mumble) um after Japan was defeated, (low mumble) uh the Japan kind of went out of China, which they had they were trying to occupy at the time.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 3.9/6; vocal-burst blend 2.0/10; 9.4s, EN.
21786_00260552 · in -25.8 dBFS · gain +5.8 dB · podcast-04630
(normal-paced, moderately variable, frequent disfluency, casual) right? Yep. There was a battle between the nationalists, (low mumble) uh, who was led by Chiang Kai shek and the communists who was led by Mao Zedong.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 4.4/6; vocal-burst blend 3.0/10; 7.7s, EN.
21786_00261672 · in -24.3 dBFS · gain +4.3 dB · podcast-04584
(normal-paced, fairly steady, some disfluency, casual) Yeah. Because they were trying to control China. And early on, Chiang Kai shek and the nationalists had a massive advantage. They had the troops, (low mumble) um, they had a little bit of support from the states.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 4.3/6; vocal-burst blend 2.4/10; 10.7s, EN.
21786_00262544 · in -25.7 dBFS · gain +5.7 dB · podcast-03522
(normal-paced, fairly steady, some disfluency, casual) The states, as we know, are trying to stop communism, so they really wanted the nationalists to win this.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 3.3/6; vocal-burst blend 2.4/10; 6.0s, EN.
21786_00263616 · in -25.4 dBFS · gain +5.4 dB · podcast-05186
Contemplation ↓  /  Angeridentity +0.64 emotion 298 %   k-AB2-k5 · #16

This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Anger up — by at least 0.25 each.

The chain starts with Anger clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.27.

At the same time Contemplation goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.11, then +0.18, then -0.05, then +0.04 — not a clean run: step 3 moves back the other way by 0.05 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.07 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.00 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.07, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 96 s · en · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.119 before conversion and 0.516 after — it rose by 0.635. Neighbour-to-neighbour the worst pair went -0.030 → 0.625. (The earlier render, with segment 1 left raw, scores 0.402 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Anger moved +0.276 in the original and +0.821 after conversion — 298 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contemplation, -0.328 became -0.361.

Quality. Mean predicted overall quality across the segments went 3.20 → 3.41 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 -0.119 → 0.516 +0.635identity cos neighbours -0.030 → 0.625d_b rescored +0.276 → +0.821d_a rescored -0.328 → -0.361d_a mined -0.329d_b mined 0.271min_cos_consec (site) -0.0045min_cos_anchor (site) -0.0663dataset podcastlang enspeaker 791454total 94.9schain gain +2.4 dBseam step 2.5 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · balanced body, some disfluency, average clarity, light breath
(contemplation, sexual lust, awe · normal-paced, normally alert, slightly relaxed, casual) No, no. Uh (ahem) you kind of touched on it in an interview I saw today that you but you've I've never couldn't find anywhere you actually talked about it. This sort of radicalism or, you know, your (ahem) uh work on behalf of people started a long time ago. I think you mentioned somewhere you were actually part of some of the freedom rides in the sixties, (low mumble) um, to Mississippi. Did I hear that?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contemplation, sexual lust, awe; style: casual, conversational; good recording, no background noise; genuineness 4.0/6; vocal-burst blend 6.1/10; 21.4s, EN.
791454_00242464 · in -20.6 dBFS · gain +0.7 dB · podcast-05076
(contemplation, doubt, interest · normal-paced, normally alert, slightly relaxed, casual) Jerry, what does the next generation of sort of activists look like to you? I mean, the the sixties doesn't happen again. You know, that sort of mentality and idealism is different now. What does it look like for you? What do you think the next round of activist lawyers looks like and what do you think their fights are, you know, for the next
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, doubt, interest; style: casual, conversational; good recording, no background noise; genuineness 3.7/6; vocal-burst blend 5.6/10; 17.4s, EN.
791454_00252833 · in -21.1 dBFS · gain +1.1 dB · podcast-06396
(contempt, bitterness, doubt · brisk, energised, neutral tension, cartoonish) Allow the person who's going to prosecute cases to write the criminal statutes. That's a legislative function. And that judiciary ought to be there to separate (low mumble) these functions. (low mumble) And we've seen a lot of encroachment, all of this deregulation. Well, I understand that bureaucracy can be a pain in the ass. (low mumble) But I want regulation of our environment. (low mumble)
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as contempt, bitterness, doubt; style: cartoonish, authoritative; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 4.2/10; 23.8s, EN.
791454_00273448 · in -19.2 dBFS · gain -0.8 dB · podcast-01476
(triumph, pride, relief · brisk, energised, neutral tension, conversational) Greed is not the only motivating factor (low mumble) that we can have. Capitalism works, but I like having fire departments. I like having police departments. I believe that everybody ought to have (low mumble) medical services (low mumble) without cost. I think that's a fundamental right. And (ahem) whether you agree with that or not, I definitely agree, (ahem) believe
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as triumph, pride, relief; style: conversational, casual; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 7.5/10; 20.4s, EN.
791454_00275828 · in -18.6 dBFS · gain -1.4 dB · podcast-06383
(anger, contempt, impatience and irritability · brisk, energised, neutral tension, cartoonish) lives (ahem) uh can and can't do and (ahem) uh with respect to the criminal law, but we ought to have them weighing in on (ahem) uh whether large corporations who most of them have become multinational,
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, fairly guarded; reads as anger, contempt, impatience and irritability; style: cartoonish, playful; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.7/10; 12.4s, EN.
791454_00278664 · in -21.8 dBFS · gain +1.8 dB · podcast-01459
Concentration ↓  /  Emotional Numbnessidentity −0.25 emotion 65 %   k-AB2-k5 · #17

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness around average — 0.55, higher than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.40.

At the same time Concentration goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.22, then +0.02, then -0.05 — not a clean run: step 4 moves back the other way by 0.05 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.75 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.59 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.75, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 48 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.648 before conversion and 0.394 after — it fell by 0.254. Neighbour-to-neighbour the worst pair went 0.729 → 0.604. (The earlier render, with segment 1 left raw, scores 0.414 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.404 in the original and +0.262 after conversion — 65 % of the delta retained. On the other named axis, Concentration, -0.342 became -0.324.

Quality. Mean predicted overall quality across the segments went 2.62 → 2.97 (+0.36) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.648 → 0.394 -0.254identity cos neighbours 0.729 → 0.604d_b rescored +0.404 → +0.262d_a rescored -0.342 → -0.324d_a mined -0.341d_b mined 0.405min_cos_consec (site) 0.5940min_cos_anchor (site) 0.7541dataset emolialang enspeaker EN_4jXpcxBn7zYtotal 46.8schain gain +1.0 dBseam step 0.6 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · slightly dark, balanced body, average recording, slightly relaxed
(concentration, pain, pride · measured, subdued, fairly steady, monologue) on the short wave infrared spectral signature. Indeed, we can observe three water absorption peak corresponding to the different wavelengths. Based on this observation, the normalized difference water index has been designed as a ratio of the difference between the near infrared and the short wave infrared on the sum of both.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, pain, pride; style: monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 4.1/10; 27.3s, EN.
EN_4jXpcxBn7zY_W000006 · in -15.3 dBFS · gain -4.7 dB · emolia-00308
(measured, normally alert, steady, formal) and turn progressively into a full green canopy spectral signature.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: formal, monologue; average recording, no background noise; genuineness 1.1/6; vocal-burst blend 0.4/10; 5.3s, EN.
EN_4jXpcxBn7zY_W000007 · in -13.6 dBFS · gain -6.5 dB · emolia-00308
(emotional numbness · measured, normally alert, fairly steady, monologue) But each of them try to improve a specific aspect, like, for instance, the savvy.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: monologue, didactic; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.2/10; 6.2s, EN.
EN_4jXpcxBn7zY_W000008 · in -14.2 dBFS · gain -5.8 dB · emolia-00308
(emotional numbness, pain · slow, normally alert, fairly steady, monologue) On the right side of the slide, an NDVI image, a light in dark.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, neutral openness; reads as emotional numbness, pain; style: monologue, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 0.6/10; 5.4s, EN.
EN_4jXpcxBn7zY_W000009 · in -15.3 dBFS · gain -4.7 dB · emolia-00308
(emotional numbness · measured, normally alert, fairly steady, monologue) Furthermore, more specific vegetation indices.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: monologue, didactic; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.4/10; 3.4s, EN.
EN_4jXpcxBn7zY_W000010 · in -13.0 dBFS · gain -7.0 dB · emolia-00308
Interest ↓  /  Emotional Numbnessidentity −0.09 emotion 118 %   k-AB2-k5 · #18

This chain comes from the two-sided rule: it only counts if both emotions move — Interest down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness around average — 0.56, higher than 56 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.34.

At the same time Interest goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.02, then +0.03, then +0.08 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 39 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.808 before conversion and 0.720 after — it fell by 0.088. Neighbour-to-neighbour the worst pair went 0.805 → 0.804. (The earlier render, with segment 1 left raw, scores 0.523 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.337 in the original and +0.400 after conversion — 118 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Interest, -0.385 became -0.608.

Quality. Mean predicted overall quality across the segments went 2.84 → 3.07 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.808 → 0.720 -0.088identity cos neighbours 0.805 → 0.804d_b rescored +0.337 → +0.400d_a rescored -0.385 → -0.608d_a mined -0.385d_b mined 0.338min_cos_consec (site) 0.8405min_cos_anchor (site) 0.8227dataset emolialang enspeaker EN_48qan7VBOo8total 37.6schain gain +2.0 dBseam step 1.5 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, measured, normally alert, slightly relaxed
(interest · steady, little disfluency, clear, didactic) Let me talk about GI science. That is geographic information science. This is the science behind the technology.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest; style: didactic, monologue; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.6/10; 7.6s, EN.
EN_48qan7VBOo8_W000043 · in -22.3 dBFS · gain +2.3 dB · emolia-01909
(concentration · steady, almost no disfluency, clear, formal) It is general knowledge that addresses the fundamental questions. These are the algorithms, data models, all those things that help us build geographic information systems, or GPS's.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: formal, monologue; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.0/10; 9.9s, EN.
EN_48qan7VBOo8_W000044 · in -20.4 dBFS · gain +0.5 dB · emolia-01909
(steady, almost no disfluency, clear, formal) The geographic technologies are those pieces of equipment used in visualization, measurement, and analysis of Earth's features.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.3/10; 6.6s, EN.
EN_48qan7VBOo8_W000045 · in -21.0 dBFS · gain +1.0 dB · emolia-01909
(steady, little disfluency, average clarity, monologue) These include things like the GIS, the GPS, the displays, the software that enable the analyst to do their job.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.5/10; 7.1s, EN.
EN_48qan7VBOo8_W000046 · in -19.9 dBFS · gain -0.1 dB · emolia-01909
(fairly steady, some disfluency, average clarity, monologue) This relationship between GI science, geographic technologies, and tradecraft might be displayed as this.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, conversational; good recording, quiet background; genuineness 1.6/6; vocal-burst blend 0.8/10; 7.2s, EN.
EN_48qan7VBOo8_W000047 · in -19.1 dBFS · gain -0.8 dB · emolia-01909
Hope Enthusiasm Optimism ↓  /  Concentrationidentity −0.02 emotion 119 %   k-AB2-k5 · #19

This chain comes from the two-sided rule: it only counts if both emotions move — Hope Enthusiasm Optimism down and Concentration up — by at least 0.25 each.

The chain starts with Concentration clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.37.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.14, then -0.12, then +0.11 — not a clean run: step 3 moves back the other way by 0.12 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.73 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.73 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.73, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 57 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.577 before conversion and 0.561 after — it fell by 0.016. Neighbour-to-neighbour the worst pair went 0.577 → 0.618. (The earlier render, with segment 1 left raw, scores 0.506 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.369 in the original and +0.439 after conversion — 119 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Hope Enthusiasm Optimism, -0.344 became -0.366.

Quality. Mean predicted overall quality across the segments went 2.66 → 2.96 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.577 → 0.561 -0.016identity cos neighbours 0.577 → 0.618d_b rescored +0.369 → +0.439d_a rescored -0.344 → -0.366d_a mined -0.344d_b mined 0.370min_cos_consec (site) 0.7342min_cos_anchor (site) 0.7342dataset emolialang enspeaker EN_hWFxh5skR1Mtotal 55.9schain gain +3.1 dBseam step 0.6 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, average recording, quiet background, normal-paced, normally alert, slightly relaxed
(hope enthusiasm optimism · average clarity, casual, monologue) How supply chains are gonna be more competitive in the future.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as hope enthusiasm optimism; style: casual, monologue; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 2.6/10; 3.0s, EN.
EN_hWFxh5skR1M_W000023 · in -15.5 dBFS · gain -4.5 dB · emolia-00858
(pride · somewhat unclear, monologue, casual) It has to do a lot with some of these recommendations that are (low mumble) in the books of Professor Sheffi, (low mumble) the Director of the Center for Transportation and Logistics at MIT, which is the boss of my boss.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride; style: monologue, casual; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 2.4/10; 11.5s, EN.
EN_hWFxh5skR1M_W000024 · in -16.3 dBFS · gain -3.7 dB · emolia-00858
(concentration, interest · somewhat unclear, monologue) And, uh, (ahem) and, and he mentioned interesting things related to this, uh, (low mumble) to this topic, how we can achieve in the supply chain more visibility. And this has to do with the fact that many of the products and services are coming from places that we haven't (low mumble) identified yet. And that means these are potential vulnerabilities.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, interest; style: monologue; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 3.9/10; 18.3s, EN.
EN_hWFxh5skR1M_W000025 · in -17.3 dBFS · gain -2.7 dB · emolia-00858
(doubt · average clarity, monologue, casual) But at least I'm trying just to reflect on, uh, (low mumble) what makes sense at this stage and why we are doing what we are doing by launching this (low mumble) sustainability course.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as doubt; style: monologue, casual; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 1.5/10; 8.3s, EN.
EN_hWFxh5skR1M_W000028 · in -17.0 dBFS · gain -3.0 dB · emolia-00858
(concentration, interest · somewhat unclear, casual, monologue) Now, when we look at, uh, (ahem) different sources, and in this case, I'm just, (low mumble) uh, citing the reference of the World Economic Forum in one of the surveys that they launched in 21, 22, which is exactly the moment of, (ahem) uh, probably the peak of the pandemic, (low mumble) uh, so far.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, interest; style: casual, monologue; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 4.1/10; 15.6s, EN.
EN_hWFxh5skR1M_W000029 · in -16.7 dBFS · gain -3.3 dB · emolia-00858
Fear ↓  /  Intoxication Altered States of Consciousnessidentity +0.11 emotion 100 %   k-AB2-k5 · #20

This chain comes from the two-sided rule: it only counts if both emotions move — Fear down and Intoxication Altered States of Consciousness up — by at least 0.25 each.

The chain starts with Intoxication Altered States of Consciousness clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.38.

At the same time Fear goes the other way, from 0.70 (higher than 70 % of clips in this corpus) to 0.32 (lower than 68 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.02, then +0.19, then +0.15, then +0.02 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 62 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.533 before conversion and 0.639 after — it rose by 0.106. Neighbour-to-neighbour the worst pair went 0.653 → 0.694. (The earlier render, with segment 1 left raw, scores 0.628 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.378 in the original and +0.379 after conversion — 100 % of the delta retained, which is essentially all of it. On the other named axis, Fear, -0.381 became -0.295.

Quality. Mean predicted overall quality across the segments went 2.96 → 3.16 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.533 → 0.639 +0.106identity cos neighbours 0.653 → 0.694d_b rescored +0.378 → +0.379d_a rescored -0.381 → -0.295d_a mined -0.381d_b mined 0.378min_cos_consec (site) 0.9099min_cos_anchor (site) 0.8899dataset emolialang enspeaker EN_lslOzUjOhjytotal 60.2schain gain +3.5 dBseam step 1.9 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a child masculine voice · fairly smooth, no background noise, normally alert, slightly relaxed, steady, no disfluency
(normal-paced, clear, fairly narrow pitch, formal) John Schoenherr – Hugo, 1935–2010
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; style: formal, casual; good recording, no background noise; no dominant emotion; genuineness 0.2/6; vocal-burst blend 3.5/10; 3.8s, EN.
EN_lslOzUjOhjy_W000059 · in -15.7 dBFS · gain -4.3 dB · emolia-00632
(normal-paced, clear, fairly narrow pitch, casual) Alex Schomburgh – 1945–1998 Vicente Segrels
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, formal; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.1/10; 5.8s, EN.
EN_lslOzUjOhjy_W000060 · in -15.5 dBFS · gain -4.5 dB · emolia-00632
(measured, clear, fairly narrow pitch, casual) CCSENF 1873–1949 Barclay Shaw
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, formal; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.2/10; 5.9s, EN.
EN_lslOzUjOhjy_W000061 · in -15.0 dBFS · gain -5.0 dB · emolia-00632
(infatuation, intoxication altered states of consciousness, longing · measured, clear, fairly narrow pitch, formal) Shusei Nagaoka – 1936–2015 John Sibak Daniel Simon Wojciech Sudmak Hajime Soriyama Simon Stalenhag Matthew Stawicki Rick Sternbach – Hugo Steve Stiles Ann Stokes Drew Struzan Ann Sudworth Arthur Soudam Daryl K.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation, intoxication altered states of consciousness, longing; style: formal, didactic; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.9/10; 29.7s, EN.
EN_lslOzUjOhjy_W000062 · in -17.8 dBFS · gain -2.2 dB · emolia-00632
(intoxication altered states of consciousness, emotional numbness · measured, crisply articulate, narrow pitch range, formal) Sean Tan' World Fantasy J.P. Target Jeff Taylor Mark Tedden Carol Thole Arthur Thompson' Adam David Thierry Timothy Truman
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is slightly cool, dark, fairly smooth, thin; crisply articulate, no disfluency, narrow pitch range, no audible breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness, emotional numbness; style: formal, didactic; average recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.0/10; 15.9s, EN.
EN_lslOzUjOhjy_W000065 · in -16.5 dBFS · gain -3.5 dB · emolia-00632