AB2: 2 chains from each of the 11 scarcest ordered emotion pairs (supply 1-5 chains each).
This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_rare-AB2-pairs.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words
are a real non-speech sound, printed where it happens.
The full generated caption for any clip is under “full caption & clip
details”. Its perceived-gender and background-noise clauses were re-rendered from
the numeric buckets, because the versions stored in the corpus index had those two
ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
The tags and captions describe the ORIGINAL clips — they are the
corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the
converted audio are the ones printed in each card's own paragraph and chips.
This chain comes from the two-sided rule: it only counts if both emotions move — Amusement down and Concentration up — by at least 0.25 each.
The chain starts with Concentration clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.25.
At the same time Amusement goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.02 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.90 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 67 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.885 before conversion and 0.888 after — it rose by 0.002. Neighbour-to-neighbour the worst pair went 0.863 → 0.845. (The earlier render, with segment 1 left raw, scores 0.683 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.248 in the original and +0.362 after conversion — 146 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Amusement, -0.251 became -0.615.
Quality. Mean predicted overall quality across the segments went 2.93 → 3.35 (+0.42) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.885 → 0.888+0.002identity cos neighbours 0.863 → 0.845d_b rescored +0.248 → +0.362d_a rescored -0.251 → -0.615d_a mined -0.251d_b mined 0.253min_cos_consec (site) 0.8967min_cos_anchor (site) 0.8914dataset podcastlang enspeaker 799798total 66.5schain gain +4.4 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, good recording, quiet background, brisk, energised, some disfluency
(amusement, hope enthusiasm optimism, interest · neutral tension, moderately variable, wide pitch range, casual)Do you have the prerequisites for it, which usually involves having a rune and being in the right spot? If the answer to that is yes, turn in your rune card, put it in the discard pile, uh (low mumble) put the event card in the discard pile, get the ordinary points, and draw on a new event card and put it face up so there's a new place people can go to get an event, right? Super simple. On your first turn, you won't have one of those probably, but (low mumble) uh they're out there. Secondly, you're gonna play one command card.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as amusement, hope enthusiasm optimism, interest; style: casual, conversational; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 6.4/10; 22.4s, EN.
799798_00838008 · in -19.5 dBFS · gain -0.5 dB · podcast-06126
(triumph, concentration, hope enthusiasm optimism ·slightly relaxed, moderately variable, wide pitch range, authoritative)Now, this is the part of the game I think is really cool mechanics. So in your hand of command cards, you have all these different actions you can do. And they're the typical actions you might see in a dudes on the map game. So one of your cards says move, one of your cards says attack, one of your cards says plunder, trade. (low mumble) Um
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as triumph, concentration, hope enthusiasm optimism; style: authoritative, dramatic; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 2.1/10; 17.7s, EN.
799798_00840240 · in -20.0 dBFS · gain -0.0 dB · podcast-06132
(concentration, contentment, hope enthusiasm optimism · slightly relaxed, fairly steady, moderate pitch range, monologue)most of those cards are the same for everybody, but every faction's got one special card that is only available to their faction. And what you're gonna do is you always you're gonna play this card face up in front of you, and the card is gonna give you a primary action, and then at the bottom of the card, there's gonna be an arrow pointing either to the left or to the right, and then have some bonus condition, and that bonus condition is gonna be based on the number of command
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as concentration, contentment, hope enthusiasm optimism; style: monologue, dramatic; good recording, quiet background; genuineness 1.4/6; vocal-burst blend 4.5/10; 26.8s, EN.
799798_00842008 · in -19.0 dBFS · gain -1.0 dB · podcast-06103
This chain comes from the two-sided rule: it only counts if both emotions move — Pleasure Ecstasy down and Elation up — by at least 0.25 each.
The chain starts with Elation at the very top of the corpus — 0.97, higher than 97 % of clips in this corpus — and works its way down to clearly present at 0.69, higher than 69 % of clips in this corpus. That is a total fall of 0.28.
At the same time Pleasure Ecstasy goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.67 (higher than 67 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are -0.09, then -0.19 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.74 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.72 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.74, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 31 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.697 before conversion and 0.675 after — it fell by 0.022. Neighbour-to-neighbour the worst pair went 0.473 → 0.389. (The earlier render, with segment 1 left raw, scores 0.499 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Elation moved -0.282 in the original and -0.236 after conversion — 84 % of the delta retained, which is most of it. On the other named axis, Pleasure Ecstasy, -0.324 became -0.439.
Quality. Mean predicted overall quality across the segments went 2.74 → 2.97 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.697 → 0.675-0.022identity cos neighbours 0.473 → 0.389d_b rescored -0.282 → -0.236d_a rescored -0.324 → -0.439d_a mined -0.324d_b mined -0.283min_cos_consec (site) 0.7233min_cos_anchor (site) 0.7410dataset emolialang enspeaker EN_B00040_S05893total 30.0schain gain +3.1 dBseam step 1.3 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · fairly smooth, balanced body, average clarity, light breath
(pleasure ecstasy, hope enthusiasm optimism, elation · brisk, energised, neutral tension, casual)I do love a laptop for its portability, but there's just something about sitting down on a desktop and diving into work that just feels more...
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as pleasure ecstasy, hope enthusiasm optimism, elation; style: casual, conversational; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 4.7/10; 6.9s, EN.
EN_B00040_S05893_W000006 · in -19.3 dBFS · gain -0.7 dB · emolia-01016
(fear, disappointment, pain· brisk, energised, neutral tension, casual)Formal. It feels more substantial. And that's why I built this disaster behind me. I wanted to have like the best system I could possibly have to sit down and tackle like more ambitious videos. I never really took advantage of this system. I think it's because I don't like sitting back here. I built this office because I wanted privacy, but I work alone in this space now. I don't need privacy.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is negative, slightly dominant, slightly guarded; reads as fear, disappointment, pain; style: casual, storytelling; average recording, quiet background; mildly explicit content; genuineness 3.6/6; vocal-burst blend 6.3/10; 20.3s, EN.
EN_B00040_S05893_W000007 · in -20.7 dBFS · gain +0.7 dB · emolia-01016
(normal-paced, normally alert, slightly relaxed, casual)So today is the day I relocate the tower setup.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 3.3/10; 3.2s, EN.
EN_B00040_S05893_W000008 · in -20.1 dBFS · gain +0.1 dB · emolia-01016
This chain comes from the two-sided rule: it only counts if both emotions move — Pleasure Ecstasy down and Elation up — by at least 0.25 each.
The chain starts with Elation at the very top of the corpus — 0.98, higher than 98 % of clips in this corpus — and works its way down to clearly present at 0.72, higher than 72 % of clips in this corpus. That is a total fall of 0.26.
At the same time Pleasure Ecstasy goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.56 (higher than 56 % of clips in this corpus), a change of -0.43. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are -0.12, then -0.14 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.31 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.28 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.31, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 38 s · pt · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.440 before conversion and 0.714 after — it rose by 0.273. Neighbour-to-neighbour the worst pair went 0.392 → 0.699. (The earlier render, with segment 1 left raw, scores 0.690 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Elation moved -0.269 in the original and -0.519 after conversion — 193 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Pleasure Ecstasy, -0.426 became -0.673.
Quality. Mean predicted overall quality across the segments went 2.47 → 3.23 (+0.76) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.440 → 0.714+0.273identity cos neighbours 0.392 → 0.699d_b rescored -0.269 → -0.519d_a rescored -0.426 → -0.673d_a mined -0.426d_b mined -0.261min_cos_consec (site) 0.2810min_cos_anchor (site) 0.3073dataset podcastlang ptspeaker 885628total 37.7schain gain +1.8 dBseam step 0.7 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · slightly cool, slightly rough, some background noise, normal-paced
(pleasure ecstasy, hope enthusiasm optimism, elation · normally alert, relaxed, fairly steady, casual)Falando isso, galera. É... (low mumble) Pra quem não tá ir por outra plataforma, procura lá no YouTube o nosso canal, o canal PolentaVerso, se você quer ver uns caras que não sabe jogar e é metido a gravar vídeo de gameplay, fazer stream jogo, vídeo de zoeira, tem ali o Benevolente lançou a nova linha Gameplay Selvagem,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is slightly cool, very dark, slightly rough, slightly thin; somewhat unclear, some disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as pleasure ecstasy, hope enthusiasm optimism, elation; style: casual; below-average recording, some background noise; genuineness 5.4/6; vocal-burst blend 10.0/10; 20.6s, PT.
885628_00051824 · in -27.0 dBFS · gain +7.0 dB · podcast-00288
(disgust, confusion, thankfulness gratitude· normally alert, neutral tension, moderately variable, casual)Você pode xingar a gente, é o que a gente gosta, né? Se você não quiser dar o seu like, se inscreve no canal pra xingar a gente. Isso
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, dark, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as disgust, confusion, thankfulness gratitude; style: casual, conversational; below-average recording, some background noise; genuineness 6.0/6; vocal-burst blend 5.8/10; 11.8s, PT.
885628_00055220 · in -27.7 dBFS · gain +7.7 dB · podcast-02629
(energised, fully relaxed, moderately variable, casual)parte mais específica que eu gostei foi a Saifa transportando o Castelvania.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, fully relaxed, moderately variable; timbre is slightly cool, very dark, slightly rough, thin; slurred, frequent disfluency, wide pitch range, audible breath; affect is positive, neutral stance, neutral openness; no dominant emotion; style: casual, playful; poor recording, some background noise; genuineness 3.6/6; vocal-burst blend 2.2/10; 5.6s, PT.
885628_00060216 · in -29.7 dBFS · gain +9.7 dB · podcast-00419
This chain comes from the two-sided rule: it only counts if both emotions move — Teasing down and Concentration up — by at least 0.25 each.
The chain starts with Concentration clearly present — 0.59, higher than 59 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.36.
At the same time Teasing goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.74 (higher than 74 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.23 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.83 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 51 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.790 before conversion and 0.888 after — it rose by 0.098. Neighbour-to-neighbour the worst pair went 0.821 → 0.900. (The earlier render, with segment 1 left raw, scores 0.735 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.360 in the original and +0.237 after conversion — 66 % of the delta retained. On the other named axis, Teasing, -0.258 became -0.082.
Quality. Mean predicted overall quality across the segments went 3.17 → 3.25 (+0.08) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.790 → 0.888+0.098identity cos neighbours 0.821 → 0.900d_b rescored +0.360 → +0.237d_a rescored -0.258 → -0.082d_a mined -0.258d_b mined 0.360min_cos_consec (site) 0.8345min_cos_anchor (site) 0.7897dataset emolialang enspeaker EN_B00008_S02659total 50.4schain gain +1.0 dBseam step 2.9 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, moderately variable, some disfluency, average clarity, wide pitch range
(teasing, malevolence malice, anger · brisk, energised, neutral tension, casual)If you don't like that, you can gargle my balls. I guess he's not gonna say that because he has a PG audience. But you get the idea. I respect what he's doing here, and I think that the real key here is not just his 200k, although his 200k is most certainly a 200k more than what I put in.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as teasing, malevolence malice, anger; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.5/6; vocal-burst blend 8.1/10; 15.1s, EN.
EN_B00008_S02659_W000000 · in -25.6 dBFS · gain +5.6 dB · emolia-00425
(anger, pride, awe· brisk, energised, neutral tension, casual)What's gonna really matter here is when people see what he did that are billionaires, or 100 millionaires, or even 10 millionaires, and decide, you know what? Here's 5 million. When they say, here's 500 million on the table, here's 200 million on the table. Because I have spoken with people that have a net worth of a billion dollars before.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as anger, pride, awe; style: casual; average recording, quiet background; mildly explicit content; genuineness 2.7/6; vocal-burst blend 6.2/10; 18.9s, EN.
EN_B00008_S02659_W000001 · in -25.3 dBFS · gain +5.3 dB · emolia-00425
(concentration, bitterness, anger ·normal-paced, normally alert, slightly relaxed, didactic)And it is not just you or I in the YouTube comments section that believe that tech companies are trying to exert too much control over our lives. It is not just people who have net worths in the three to five figures that believe that tech companies are becoming bullies and that they need to have their power checked.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration, bitterness, anger; style: didactic, casual; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.4/10; 16.8s, EN.
EN_B00008_S02659_W000002 · in -27.4 dBFS · gain +7.4 dB · emolia-00425
This chain comes from the two-sided rule: it only counts if both emotions move — Teasing down and Concentration up — by at least 0.25 each.
The chain starts with Concentration around average — 0.50, right about the corpus median — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.33.
At the same time Teasing goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.12, then +0.20 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 37 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.916 before conversion and 0.856 after — it fell by 0.060. Neighbour-to-neighbour the worst pair went 0.916 → 0.860. (The earlier render, with segment 1 left raw, scores 0.809 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.328 in the original and +0.356 after conversion — 109 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Teasing, -0.261 became -0.626.
Quality. Mean predicted overall quality across the segments went 3.13 → 3.31 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.916 → 0.856-0.060identity cos neighbours 0.916 → 0.860d_b rescored +0.328 → +0.356d_a rescored -0.261 → -0.626d_a mined -0.261d_b mined 0.326min_cos_consec (site) 0.9256min_cos_anchor (site) 0.9256dataset emolialang enspeaker EN_3OgzBKWK6qAtotal 36.3schain gain +2.6 dBseam step 0.7 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, balanced body, average recording, quiet background, slightly relaxed, fairly steady, average clarity, moderate pitch range
(teasing, infatuation, intoxication altered states of consciousness · measured, normally alert, frequent disfluency, didactic)Now if you had Anthony, then one of your bench players would have come on. This is the order they would have appeared on your bench. Of course you only had three of these on your bench. And they scored one, seven, one, seven, two, six, one and zero.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as teasing, infatuation, intoxication altered states of consciousness; style: didactic, monologue; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 1.2/10; 14.1s, EN.
EN_3OgzBKWK6qA_W000012 · in -18.4 dBFS · gain -1.6 dB · emolia-00723
(sourness, fear, pain·normal-paced, normally alert, some disfluency, conversational)So by my reckoning, the global average was 53. The worst you could have got, if you were as unlucky as possible, was 26 of this system. The average was 45. The best was 70.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, fear, pain; style: conversational, casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.8/10; 11.2s, EN.
EN_3OgzBKWK6qA_W000013 · in -17.5 dBFS · gain -2.5 dB · emolia-00723
(measured, subdued, frequent disfluency, casual)And certainly everyone I see who's doing this is between the 5 and 10% mark, so should be fine for finishing top 5% I think, as things stand. This is to remind me to say
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 1.9/10; 11.4s, EN.
EN_3OgzBKWK6qA_W000014 · in -19.0 dBFS · gain -1.0 dB · emolia-00723
This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Amusement up — by at least 0.25 each.
The chain starts with Amusement clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.27.
At the same time Concentration goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.11, then +0.16 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.87 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 40 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.750 before conversion and 0.703 after — it fell by 0.047. Neighbour-to-neighbour the worst pair went 0.762 → 0.818. (The earlier render, with segment 1 left raw, scores 0.600 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Amusement moved +0.271 in the original and +0.613 after conversion — 226 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.282 became -0.204.
Quality. Mean predicted overall quality across the segments went 2.95 → 3.17 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.750 → 0.703-0.047identity cos neighbours 0.762 → 0.818d_b rescored +0.271 → +0.613d_a rescored -0.282 → -0.204d_a mined -0.283d_b mined 0.271min_cos_consec (site) 0.8701min_cos_anchor (site) 0.9132dataset emolialang enspeaker EN_B00012_S00311total 39.2schain gain +1.8 dBseam step 1.4 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, slightly rough, balanced body, quiet background, slightly relaxed, fairly steady, average clarity
(concentration, doubt, sexual lust · measured, subdued, frequent disfluency, didactic)Which is equal to the angle of the wedge that I've attached to my jig. Now if you're worried about all these dimensions, if you choose this method of construction, the jig will be shown in the measured drawings. I'm gonna make the mortise using this quarter inch spiral bit and this guide collar. It slips into the pre-made slot and I'll just plunge the bit into the work.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as concentration, doubt, sexual lust; style: didactic, monologue; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 0.3/10; 24.7s, EN.
EN_B00012_S00311_W000070 · in -21.5 dBFS · gain +1.5 dB · emolia-00483
(sexual lust ·slow, normally alert, frequent disfluency, monologue)Okay, now I'm cutting out the slats with my jigsaw. And I laid out each piece with a paper pattern to show the form.
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, neutral openness; reads as sexual lust; style: monologue, didactic; good recording, quiet background; genuineness 1.8/6; vocal-burst blend 0.6/10; 8.4s, EN.
EN_B00012_S00311_W000071 · in -20.9 dBFS · gain +0.9 dB · emolia-00483
(amusement, teasing, sexual lust ·measured, normally alert, some disfluency, monologue)Now all I have to do is smooth the edges with the drum sander, bevel them a little bit, and fit them in the mortises.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as amusement, teasing, sexual lust; style: monologue, didactic; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 0.6/10; 6.5s, EN.
EN_B00012_S00311_W000072 · in -20.1 dBFS · gain +0.1 dB · emolia-00483
This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Amusement up — by at least 0.25 each.
The chain starts with Amusement clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.26.
At the same time Concentration goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.07, then +0.01, then +0.18 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.90 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 105 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.910 before conversion and 0.858 after — it fell by 0.051. Neighbour-to-neighbour the worst pair went 0.927 → 0.858. (The earlier render, with segment 1 left raw, scores 0.607 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Amusement moved +0.260 in the original and +0.515 after conversion — 198 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.417 became -0.341.
Quality. Mean predicted overall quality across the segments went 2.71 → 3.19 (+0.49) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.910 → 0.858-0.051identity cos neighbours 0.927 → 0.858d_b rescored +0.260 → +0.515d_a rescored -0.417 → -0.341d_a mined -0.413d_b mined 0.260min_cos_consec (site) 0.8999min_cos_anchor (site) 0.8909dataset podcastlang enspeaker 859092total 104.5schain gain +3.7 dBseam step 1.0 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-bright, slightly rough, balanced body, average recording, quiet background, neutral tension, moderately variable, wide pitch range
(concentration, bitterness, triumph · measured, normally alert, frequent disfluency, didactic)Pretty strong language, right? Very strong language. The message to the contemporary church couldn't be clearer. The United States, among other things, and most of the contemporary world is filled with churches that are lukewarm, complacent, apathetic. You know what we would call it? Comfortable.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, fairly guarded; reads as concentration, bitterness, triumph; style: didactic, authoritative; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.1/10; 24.3s, EN.
859092_00127096 · in -15.1 dBFS · gain -4.9 dB · podcast-05007
(sourness, disgust, malevolence malice·brisk, energised, some disfluency, dramatic)If you're comfortable in your faith with Jesus Christ, you're lukewarm, baby. Don't get comfortable. You know what real comfortable is? It's called horizontal room temperature. That's comfortable. That's how Jesus views this. He can't stand it. Now, that's not how the church saw itself. What was the church's self image? Man, these people had quite an image, verse 17. Because you say, I am rich and have become wealthy and have need of nothing.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; clear, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as sourness, disgust, malevolence malice; style: dramatic, cartoonish; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 2.7/10; 29.3s, EN.
859092_00129528 · in -12.9 dBFS · gain -7.1 dB · podcast-05009
(interest, jealousy and envy, triumph· brisk, energised, some disfluency, conversational)Wow. Here's the principle. When you trust in wealth, you lie to yourself. When you trust in wealth, you lie to yourself. I was going to put down when you trust in your wealth, and I thought, no, no, you can be broke and trust in wealth. Broke people trust in wealth all the time. You're still lying to yourself because you believe if you just had enough wealth, you wouldn't need anything else. That's what this group says. Some people believe that if you have enough money, you literally don't. I mean, who needs Jesus if you have enough cash?
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as interest, jealousy and envy, triumph; style: conversational, casual; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 7.3/10; 27.9s, EN.
859092_00132452 · in -12.7 dBFS · gain -7.3 dB · podcast-05004
(amusement, elation, relief· brisk, energised, some disfluency, didactic)Right? Until the doctor calls you up and says, by the way, your terminal. (ahem) Then you start getting serious about something beyond this life, right? It's interesting, this group was very self-sufficient because they said, Well, I have become wealthy. Sounds like recent wealth acquisition, and this was a very diligent, hardworking town. So this group actually earned the money.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as amusement, elation, relief; style: didactic, casual; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 3.3/10; 23.5s, EN.
859092_00135240 · in -13.2 dBFS · gain -6.8 dB · podcast-05023
This chain comes from the two-sided rule: it only counts if both emotions move — Teasing down and Contempt up — by at least 0.25 each.
The chain starts with Contempt clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.37.
At the same time Teasing goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.21 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.61 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.52 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.61, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 19 s · zh · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.536 before conversion and 0.744 after — it rose by 0.208. Neighbour-to-neighbour the worst pair went 0.476 → 0.712. (The earlier render, with segment 1 left raw, scores 0.554 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contempt moved +0.374 in the original and +0.603 after conversion — 161 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Teasing, -0.252 became +0.102.
Quality. Mean predicted overall quality across the segments went 2.71 → 3.10 (+0.40) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.536 → 0.744+0.208identity cos neighbours 0.476 → 0.712d_b rescored +0.374 → +0.603d_a rescored -0.252 → +0.102d_a mined -0.252d_b mined 0.373min_cos_consec (site) 0.5181min_cos_anchor (site) 0.6067dataset emolialang zhspeaker ZH_B00049_S07542total 18.4schain gain +1.2 dBseam step 0.1 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · thin, fast
A child feminine voice; delivery is energised, fast, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; very clear, almost no disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, guarded; reads as confusion, distress; style: cartoonish, storytelling; average recording, no background noise; genuineness 2.0/6; vocal-burst blend 4.0/10; 8.9s, ZH.
ZH_B00049_S07542_W000003 · in -17.3 dBFS · gain -2.7 dB · emolia-03764
(contempt, impatience and irritability, helplessness·highly aroused, very tense, volatile, cartoonish)我真为老爸捏一把汗呢,老爸,你可真要挺住啊。
full caption & clip details
A child strongly feminine voice; delivery is highly aroused, fast, very tense, volatile; timbre is cool, bright, very rough, thin; slurred, almost no disfluency, very wide pitch range, normal breath; affect is elated, slightly dominant, guarded; reads as contempt, impatience and irritability, helplessness; style: cartoonish, storytelling; below-average recording, some background noise; genuineness 1.9/6; vocal-burst blend 4.0/10; 5.9s, ZH.
ZH_B00049_S07542_W000004 · in -15.1 dBFS · gain -4.9 dB · emolia-03764
This chain comes from the two-sided rule: it only counts if both emotions move — Teasing down and Contempt up — by at least 0.25 each.
The chain starts with Contempt clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.30.
At the same time Teasing goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.13 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.55 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.72 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.55, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 52 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.533 before conversion and 0.621 after — it rose by 0.088. Neighbour-to-neighbour the worst pair went 0.683 → 0.732. (The earlier render, with segment 1 left raw, scores 0.531 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contempt moved +0.300 in the original and +0.303 after conversion — 101 % of the delta retained, which is essentially all of it. On the other named axis, Teasing, -0.634 became -0.633.
Quality. Mean predicted overall quality across the segments went 2.66 → 3.10 (+0.44) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.533 → 0.621+0.088identity cos neighbours 0.683 → 0.732d_b rescored +0.300 → +0.303d_a rescored -0.634 → -0.633d_a mined -0.267d_b mined 0.301min_cos_consec (site) 0.7245min_cos_anchor (site) 0.5468dataset podcastlang enspeaker 586941total 51.6schain gain +1.8 dBseam step 1.7 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, normally alert, neutral tension, some disfluency, light breath
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, jealousy and envy, embarrassment; style: conversational, monologue; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 9.8/10; 17.2s, EN.
586941_00039062 · in -24.7 dBFS · gain +4.7 dB · podcast-04974
(contempt·fast, moderately variable, average clarity, casual)Again. Go gajita pilos in the goasi guakasi internet, the gochasi, a papan, kaunik matic mission and
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contempt; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 6.5/10; 7.9s, EN.
586941_00042104 · in -23.1 dBFS · gain +3.1 dB · podcast-03169
This chain comes from the two-sided rule: it only counts if both emotions move — Sourness down and Teasing up — by at least 0.25 each.
The chain starts with Teasing clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.25.
At the same time Sourness goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.07 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.59 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.59 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.59, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 32 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.455 before conversion and 0.625 after — it rose by 0.170. Neighbour-to-neighbour the worst pair went 0.455 → 0.625. (The earlier render, with segment 1 left raw, scores 0.478 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Teasing moved +0.253 in the original and +0.101 after conversion — 40 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Sourness, -0.279 became -0.354.
Quality. Mean predicted overall quality across the segments went 2.77 → 3.05 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.455 → 0.625+0.170identity cos neighbours 0.455 → 0.625d_b rescored +0.253 → +0.101d_a rescored -0.279 → -0.354d_a mined -0.289d_b mined 0.254min_cos_consec (site) 0.5856min_cos_anchor (site) 0.5856dataset emolialang enspeaker EN_B00007_S09239total 31.6schain gain +2.5 dBseam step 1.4 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, average recording, normally alert, some disfluency, average clarity
(sourness, elation, interest · brisk, neutral tension, moderately variable, conversational)Yeaah, yeaah, shellfish. And then all of that, then you'll have the pasta, so then you'll have the pasta with the oil and the vinegar, and then probably like (ahem) sardines would be mixed into one of the pastas, because sardines will be part of it, and then you'll have the, oh, not sardines, anchovies.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as sourness, elation, interest; style: conversational, casual; average recording, some background noise; mildly explicit content; genuineness 5.6/6; vocal-burst blend 9.2/10; 18.6s, EN.
EN_B00007_S09239_W000096 · in -25.0 dBFS · gain +5.0 dB · emolia-00407
(sexual lust, amusement, disgust·normal-paced, neutral tension, moderately variable, casual)I bet that house fucking stinks. Yeah, and then something else. So yeah. Yeah.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, very rough, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as sexual lust, amusement, disgust; style: casual, conversational; average recording, quiet background; explicit content; genuineness 5.9/6; vocal-burst blend 0.6/10; 4.2s, EN.
EN_B00007_S09239_W000097 · in -23.9 dBFS · gain +4.0 dB · emolia-00407
(teasing, confusion, infatuation· normal-paced, slightly relaxed, fairly steady, conversational)So what's your, what's the, you're on the PlayStation. What's your next go-to or one of your that you haven't talked about maybe that you've been playing?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as teasing, confusion, infatuation; style: conversational, casual; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 1.5/10; 9.2s, EN.
EN_B00007_S09239_W000098 · in -27.5 dBFS · gain +7.5 dB · emolia-00407
This chain comes from the two-sided rule: it only counts if both emotions move — Sourness down and Teasing up — by at least 0.25 each.
The chain starts with Teasing clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.26.
At the same time Sourness goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.11, then +0.10, then +0.05 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.35 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.30 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.35, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 54 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.199 before conversion and 0.670 after — it rose by 0.471. Neighbour-to-neighbour the worst pair went 0.199 → 0.626. (The earlier render, with segment 1 left raw, scores 0.634 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Teasing moved +0.258 in the original and +0.079 after conversion — 31 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Sourness, -0.341 became -0.031.
Quality. Mean predicted overall quality across the segments went 2.91 → 3.06 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.199 → 0.670+0.471identity cos neighbours 0.199 → 0.626d_b rescored +0.258 → +0.079d_a rescored -0.341 → -0.031d_a mined -0.342d_b mined 0.256min_cos_consec (site) 0.2999min_cos_anchor (site) 0.3483dataset podcastlang enspeaker 24341total 53.0schain gain +4.7 dBseam step 1.9 dBcrossfades 100/150/100 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average clarity
(sourness, disappointment, affection · normal-paced, normally alert, neutral tension, casual)none, right? So I actually pity I mean, I feel sorry for her more than I hate her, but I also hate her because of how she's so like she talks and talks and talks and talks. She's such a she's the gossip, right? She's the gossip and rumor mill person of the book, and it's so frustrating.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as sourness, disappointment, affection; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 3.8/6; vocal-burst blend 6.6/10; 18.9s, EN.
24341_00547328 · in -28.2 dBFS · gain +8.2 dB · podcast-06197
(disgust, bitterness, sourness ·measured, normally alert, slightly relaxed, casual)Yeah, and it's gossip for the pure sake of being the person who knows things.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as disgust, bitterness, sourness; style: casual, conversational; good recording, no background noise; explicit content; genuineness 3.2/6; vocal-burst blend 2.5/10; 4.2s, EN.
24341_00549248 · in -32.1 dBFS · gain +12.1 dB · podcast-00833
(doubt, intoxication altered states of consciousness, amusement·normal-paced, normally alert, relaxed, casual)Yeah. And you know what? (ahem) Um Maybe this is a a catty thing, but it feels like and I mean I'm not an expert on any of this, so take it with a grain of salt, but the way that Rand writes about Lillian feels like the way a woman would talk about another woman they don't like.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as doubt, intoxication altered states of consciousness, amusement; style: casual, monologue; good recording, no background noise; genuineness 4.6/6; vocal-burst blend 6.9/10; 19.2s, EN.
24341_00549680 · in -28.6 dBFS · gain +8.6 dB · podcast-00840
(teasing, embarrassment, amusement · normal-paced, energised, relaxed, casual)what I mean? Yeah, I know exactly what you mean. And I think you're right. Apologies all around to all listeners out there. Okay. The most (ahem) of all of the villains in the book,
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as teasing, embarrassment, amusement; style: casual, conversational; average recording, some background noise; mildly explicit content; genuineness 6.0/6; vocal-burst blend 2.5/10; 11.1s, EN.
24341_00551752 · in -19.0 dBFS · gain -1.0 dB · podcast-03532
This chain comes from the two-sided rule: it only counts if both emotions move — Teasing down and Pleasure Ecstasy up — by at least 0.25 each.
The chain starts with Pleasure Ecstasy clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.30.
At the same time Teasing goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.74 (higher than 74 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.06, then +0.01 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.44 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.31 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.44, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 25 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.403 before conversion and 0.664 after — it rose by 0.261. Neighbour-to-neighbour the worst pair went 0.367 → 0.618. (The earlier render, with segment 1 left raw, scores 0.498 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pleasure Ecstasy moved +0.285 in the original and +0.077 after conversion — 27 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Teasing, -0.254 became -0.631.
Quality. Mean predicted overall quality across the segments went 2.57 → 2.91 (+0.33) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.403 → 0.664+0.261identity cos neighbours 0.367 → 0.618d_b rescored +0.285 → +0.077d_a rescored -0.254 → -0.631d_a mined -0.251d_b mined 0.299min_cos_consec (site) 0.3135min_cos_anchor (site) 0.4401dataset podcastlang enspeaker 633786total 24.2schain gain -1.1 dBseam step 2.3 dBcrossfades 150/100/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, moderately variable, light breath
(teasing, embarrassment, confusion · normally alert, neutral tension, some disfluency, casual)When he was four. Okay. I'm (ahem) kidding. But no, he's in (ahem) um, if you look, there's a (low mumble) uh a brewing book that he's actually in and everything.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as teasing, embarrassment, confusion; style: casual, conversational; average recording, some background noise; mildly explicit content; genuineness 5.9/6; vocal-burst blend 3.8/10; 9.6s, EN.
633786_00173192 · in -19.8 dBFS · gain -0.2 dB · podcast-02350
(amusement, teasing, intoxication altered states of consciousness·energised, neutral tension, some disfluency, casual)I'm putting you on the wheel for that. Color printers were more expensive. (childlike giggle) Color printers are expensive now. Well,
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as amusement, teasing, intoxication altered states of consciousness; style: casual, playful; average recording, some background noise; mildly explicit content; genuineness 6.0/6; vocal-burst blend 2.3/10; 7.2s, EN.
633786_00175008 · in -17.5 dBFS · gain -2.5 dB · podcast-00171
(pleasure ecstasy, longing, shame·normally alert, slightly relaxed, frequent disfluency, casual)(low mumble) Um I like this better than the other one. I still don't like it very much.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as pleasure ecstasy, longing, shame; style: casual, conversational; average recording, no background noise; genuineness 4.4/6; vocal-burst blend 4.6/10; 3.7s, EN.
633786_00176112 · in -21.3 dBFS · gain +1.3 dB · podcast-02336
(pleasure ecstasy, elation, astonishment surprise· normally alert, slightly relaxed, frequent disfluency, conversational)right. I'm gonna go with a seven. I actually truly enjoyed that.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as pleasure ecstasy, elation, astonishment surprise; style: conversational, casual; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 2.0/10; 4.1s, EN.
633786_00176720 · in -25.5 dBFS · gain +5.5 dB · podcast-02354
This chain comes from the two-sided rule: it only counts if both emotions move — Teasing down and Pleasure Ecstasy up — by at least 0.25 each.
The chain starts with Pleasure Ecstasy clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.32.
At the same time Teasing goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.09 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.69 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.68 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.69, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 22 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.719 before conversion and 0.730 after — it rose by 0.011. Neighbour-to-neighbour the worst pair went 0.586 → 0.703. (The earlier render, with segment 1 left raw, scores 0.638 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pleasure Ecstasy moved +0.318 in the original and +0.677 after conversion — 213 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Teasing, -0.258 became -0.128.
Quality. Mean predicted overall quality across the segments went 2.72 → 3.00 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.719 → 0.730+0.011identity cos neighbours 0.586 → 0.703d_b rescored +0.318 → +0.677d_a rescored -0.258 → -0.128d_a mined -0.256d_b mined 0.318min_cos_consec (site) 0.6824min_cos_anchor (site) 0.6904dataset podcastlang enspeaker 65232total 21.5schain gain +3.2 dBseam step 3.4 dBcrossfades 150/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert
(teasing, sourness, embarrassment · normal-paced, neutral tension, moderately variable, casual)it's it's it's tough when you're not the best Hernan Gomez, you know, when you're not even the best (ahem) Hernan. Uh I think Wancho is
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as teasing, sourness, embarrassment; style: casual, playful; average recording, quiet background; mildly explicit content; genuineness 5.8/6; vocal-burst blend 1.7/10; 7.8s, EN.
65232_00067736 · in -26.3 dBFS · gain +6.3 dB · podcast-00316
(teasing, elation· normal-paced, neutral tension, fairly steady, casual)currently at top, not only because he has the better name, but I do enjoy the Charlotte (low mumble) uh Hornets (ahem) um signing (ahem) uh spending on a max contract to get Gordon Hayward just to move from ninth to 10th in the conference.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as teasing, elation; style: casual, conversational; good recording, quiet background; genuineness 4.9/6; vocal-burst blend 6.6/10; 10.6s, EN.
65232_00068508 · in -25.6 dBFS · gain +5.6 dB · podcast-02110
(pleasure ecstasy, elation, embarrassment·measured, fully relaxed, fairly steady, casual)(low mumble) Um really excited for them to do that, honestly. Uh (low mumble)
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, fully relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as pleasure ecstasy, elation, embarrassment; style: casual, conversational; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 1.1/10; 3.5s, EN.
65232_00069568 · in -30.2 dBFS · gain +10.2 dB · podcast-04390
This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Teasing up — by at least 0.25 each.
The chain starts with Teasing clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.27.
At the same time Contemplation goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.45. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.21, then -0.04, then +0.09 — not a clean run: step 2 moves back the other way by 0.04 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 94 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.895 before conversion and 0.909 after — it rose by 0.014. Neighbour-to-neighbour the worst pair went 0.923 → 0.873. (The earlier render, with segment 1 left raw, scores 0.529 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Teasing moved +0.266 in the original and +0.113 after conversion — 42 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Contemplation, -0.443 became -0.323.
Quality. Mean predicted overall quality across the segments went 2.50 → 3.22 (+0.72) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.895 → 0.909+0.014identity cos neighbours 0.923 → 0.873d_b rescored +0.266 → +0.113d_a rescored -0.443 → -0.323d_a mined -0.448d_b mined 0.267min_cos_consec (site) 0.9244min_cos_anchor (site) 0.9144dataset podcastlang enspeaker 846612total 93.4schain gain +2.9 dBseam step 0.1 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · dark, below-average recording, relaxed, frequent disfluency
(contemplation, relief, malevolence malice · measured, very low-energy, fairly steady, casual)Let it do and keep yourself in thought. You know what I'm saying? Try to work on yourself physically, spiritually, mentally, psychologically, you know. Get your ass up in the gym to try and get them in conference pumping so you can keep that inner happiness smaller within that momentum. You feel me? (low mumble) Um
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is warm, dark, rough, slightly thin; somewhat unclear, frequent disfluency, narrow pitch range, normal breath; affect is mildly negative, slightly dominant, fairly guarded; reads as contemplation, relief, malevolence malice; style: casual, monologue; below-average recording, quiet background; genuineness 4.4/6; vocal-burst blend 6.8/10; 21.0s, EN.
846612_00120152 · in -23.9 dBFS · gain +3.9 dB · podcast-00582
(sourness, thankfulness gratitude, contentment· measured, very low-energy, fairly steady, storytelling)what else am I supposed to say? You know, fed people thoroughly. Don't be afraid to talk to people, because you know people can't lead to money, you know what I'm saying? Because you need to learn how to probably speak to people and say they body language, and also with the temperature of the room,
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is warm, dark, rough, very full; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is negative, slightly dominant, fairly guarded; reads as sourness, thankfulness gratitude, contentment; style: storytelling, monologue; below-average recording, some background noise; genuineness 2.5/6; vocal-burst blend 2.7/10; 19.5s, EN.
846612_00122244 · in -24.4 dBFS · gain +4.4 dB · podcast-00576
(sourness, impatience and irritability, bitterness·slow, very low-energy, fairly steady, storytelling)right? Because unfortunately, if you do not read the temperature of the room, and you don't set the cues, actually you bounce the fuck up out of here, if you don't embrace getting rejected, if you don't embrace the no's that you hear all the time, if you don't embrace those things, you know, if you don't embrace, you know, a woman or another man, you
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is cool, dark, rough, very full; slurred, frequent disfluency, wide pitch range, audible breath; affect is negative, slightly dominant, fairly guarded; reads as sourness, impatience and irritability, bitterness; style: storytelling, whispered; below-average recording, quiet background; genuineness 3.1/6; vocal-burst blend 4.6/10; 24.8s, EN.
846612_00124192 · in -27.2 dBFS · gain +7.2 dB · podcast-00569
(teasing, bitterness, anger·normal-paced, normally alert, moderately variable, storytelling)know, saying they got significant other, right? If you can't handheld those types of things, right. You can't handle them, you want to try and be persistent, you wanna go try and go crazy, you know what I'm saying? God forbid, you gotta like Myron Gaines, the freshman CEO of Fresnel Fit, you know, again to tend to be me too by some 19-year-old arrogant chick. You know what I'm saying? As I discussed in like the last podcast.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, dark, slightly rough, balanced body; slurred, frequent disfluency, wide pitch range, no audible breath; affect is positive, slightly dominant, fairly guarded; reads as teasing, bitterness, anger; style: storytelling, casual; below-average recording, some background noise; genuineness 3.4/6; vocal-burst blend 5.5/10; 28.7s, EN.
846612_00126665 · in -24.1 dBFS · gain +4.1 dB · podcast-00565
This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Teasing up — by at least 0.25 each.
The chain starts with Teasing clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.26.
At the same time Contemplation goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.02, then +0.23 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.91 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 74 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.907 before conversion and 0.901 after — it fell by 0.006. Neighbour-to-neighbour the worst pair went 0.907 → 0.912. (The earlier render, with segment 1 left raw, scores 0.644 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Teasing moved +0.259 in the original and +0.139 after conversion — 54 % of the delta retained. On the other named axis, Contemplation, -0.325 became -0.396.
Quality. Mean predicted overall quality across the segments went 3.25 → 3.36 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.907 → 0.901-0.006identity cos neighbours 0.907 → 0.912d_b rescored +0.259 → +0.139d_a rescored -0.325 → -0.396d_a mined -0.320d_b mined 0.259min_cos_consec (site) 0.9059min_cos_anchor (site) 0.8986dataset podcastlang enspeaker 130501total 73.6schain gain +3.3 dBseam step 3.5 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, fairly steady, light breath
(contemplation, shame, jealousy and envy · measured, subdued, neutral tension, casual)Yeah. Yeah. I have multiple times on this show mentioned that I don't think it should be its own vehicle. I love it. They've done a a good job with it, but it shouldn't be its own car. Part of part of the great thing about Subaru was, and can be to an extent, that no matter what you got from them, you could get something out of it.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as contemplation, shame, jealousy and envy; style: casual, conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 7.3/10; 22.6s, EN.
130501_00341632 · in -28.1 dBFS · gain +8.1 dB · podcast-06420
(interest, elation, disgust·brisk, normally alert, slightly relaxed, casual)Performance wise. And while that's probably that's still true with the WRX, with the other models, it's really not. There's not a ton of standard plug and play aftermarket for stuff like the Impreza or the Forester anymore. And I think part of that is because they took the WRX. Now it wasn't always a WRX, right? But in Japan, it was for the Forester. It was an STI. And the Impreza got the trim as well. And they put it on its own thing and called it a day. I think you miss out on a couple very, very good cars doing it that way.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as interest, elation, disgust; style: casual, monologue; good recording, quiet background; genuineness 2.6/6; vocal-burst blend 9.8/10; 24.9s, EN.
130501_00343884 · in -27.8 dBFS · gain +7.8 dB · podcast-06428
(teasing, triumph, pride·normal-paced, normally alert, slightly relaxed, casual)And I think it's just easier as a brand for Super just to go, well, here's the WRX. If you want that, there it is. All right, but you know, what if I want what if I want a an Impreza wagon with a turbocharged four cylinder and a five-banger in it? Why why do I gotta go buy this 40 something thousand dollar sedan? So I know, and it's (ahem) uh listen, I like the car. I do. You see them on the road, people have done some cool stuff with them. You see them in that blue with the gold rims and spoilers, they look good.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as teasing, triumph, pride; style: casual, conversational; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 10.0/10; 26.5s, EN.
130501_00346376 · in -28.3 dBFS · gain +8.3 dB · podcast-06407
This chain comes from the two-sided rule: it only counts if both emotions move — Amusement down and Teasing up — by at least 0.25 each.
The chain starts with Teasing clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.27.
At the same time Amusement goes the other way, from 0.74 (higher than 74 % of clips in this corpus) to 0.99 (higher than 99 % of clips in this corpus), a change of +0.25. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.14 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.11 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.11 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.11, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 19 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.032 before conversion and 0.413 after — it rose by 0.381. Neighbour-to-neighbour the worst pair went 0.032 → 0.455. (The earlier render, with segment 1 left raw, scores 0.304 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Teasing moved +0.266 in the original and +0.185 after conversion — 69 % of the delta retained. On the other named axis, Amusement, +0.252 became +0.166.
Quality. Mean predicted overall quality across the segments went 2.40 → 2.84 (+0.44) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.032 → 0.413+0.381identity cos neighbours 0.032 → 0.455d_b rescored +0.266 → +0.185d_a rescored +0.252 → +0.166d_a mined 0.252d_b mined 0.267min_cos_consec (site) 0.1114min_cos_anchor (site) 0.1114dataset emolialang enspeaker EN_iAClv9RwJ6Mtotal 18.0schain gain +2.1 dBseam step 1.7 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · normal-paced, some disfluency
(normally alert, slightly relaxed, fairly steady, casual)As soon as we take those stems back, let's go look for...
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 1.0/10; 3.0s, EN.
EN_iAClv9RwJ6M_W000030 · in -19.1 dBFS · gain -0.9 dB · emolia-00859
(astonishment surprise, intoxication altered states of consciousness, sexual lust· normally alert, relaxed, moderately variable, casual)Yeah, I think we should do this. Oh, I found an ant cave. Oh! Fucking gnats. Let's get this shit out of me.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is slightly cool, dark, fairly smooth, slightly thin; slurred, some disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as astonishment surprise, intoxication altered states of consciousness, sexual lust; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 5.5/6; vocal-burst blend 0.7/10; 7.4s, EN.
EN_iAClv9RwJ6M_W000031 · in -20.3 dBFS · gain +0.3 dB · emolia-00859
(teasing, amusement, infatuation·energised, relaxed, moderately variable, casual)Into a pot, you just literally (ahem) dropped them all down the Xanthill. (chuckle) (ahem) You were supposed to bring them back, what are you doing?
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as teasing, amusement, infatuation; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 5.8/6; vocal-burst blend 0.0/10; 7.9s, EN.
EN_iAClv9RwJ6M_W000034 · in -20.4 dBFS · gain +0.4 dB · emolia-00859
This chain comes from the two-sided rule: it only counts if both emotions move — Amusement down and Teasing up — by at least 0.25 each.
The chain starts with Teasing clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.27.
At the same time Amusement goes the other way, from 0.74 (higher than 74 % of clips in this corpus) to 0.99 (higher than 99 % of clips in this corpus), a change of +0.25. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.18, then +0.09 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.32 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.45 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.32, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 40 s · es · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.242 before conversion and 0.339 after — it rose by 0.096. Neighbour-to-neighbour the worst pair went 0.242 → 0.339. (The earlier render, with segment 1 left raw, scores 0.406 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Teasing moved +0.265 in the original and +0.011 after conversion — 4 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Amusement, +0.252 became +0.085.
Quality. Mean predicted overall quality across the segments went 2.36 → 2.75 (+0.39) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.242 → 0.339+0.096identity cos neighbours 0.242 → 0.339d_b rescored +0.265 → +0.011d_a rescored +0.252 → +0.085d_a mined 0.253d_b mined 0.270min_cos_consec (site) 0.4474min_cos_anchor (site) 0.3207dataset podcastlang esspeaker 516660total 39.1schain gain +2.0 dBseam step 1.9 dBcrossfades 100/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · slightly rough, normal-paced, energised, moderately variable, wide pitch range
(slightly relaxed, some disfluency, somewhat unclear, casual)still on a solar letrista tonta, okay? Well,
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; no dominant emotion; style: casual, conversational; poor recording, some background noise; genuineness 5.3/6; vocal-burst blend 3.6/10; 3.2s, ES.
516660_00196744 · in -23.7 dBFS · gain +3.7 dB · podcast-02442
(pleasure ecstasy, elation, affection·neutral tension, frequent disfluency, somewhat unclear, casual)Thank you. (low mumble) And then we continue to have an (surprised gasp) (childlike giggle) announcement (surprised gasp) special.
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as pleasure ecstasy, elation, affection; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 6.5/10; 23.3s, ES.
516660_00202288 · in -25.1 dBFS · gain +5.1 dB · podcast-00737
(teasing, embarrassment, amusement· neutral tension, some disfluency, slurred, casual)se hace como el se hace el inteligente no sí ok que las oraciones que voy a decir son entradas del Pokedex para que sea más cool la cosa ok
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, dark, slightly rough, slightly thin; slurred, some disfluency, wide pitch range, heavy breath; affect is positive, slightly submissive, neutral openness; reads as teasing, embarrassment, amusement; style: casual, playful; below-average recording, some background noise; genuineness 5.4/6; vocal-burst blend 5.7/10; 12.9s, ES.
516660_00213688 · in -26.9 dBFS · gain +6.9 dB · podcast-00180
This chain comes from the two-sided rule: it only counts if both emotions move — Teasing down and Amusement up — by at least 0.25 each.
The chain starts with Amusement clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.27.
At the same time Teasing goes the other way, from 0.73 (higher than 73 % of clips in this corpus) to 0.99 (higher than 99 % of clips in this corpus), a change of +0.26. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.14 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.74 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.74 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.74, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 33 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.302 before conversion and 0.496 after — it rose by 0.195. Neighbour-to-neighbour the worst pair went 0.302 → 0.579. (The earlier render, with segment 1 left raw, scores 0.293 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Amusement moved +0.267 in the original and +0.194 after conversion — 73 % of the delta retained, which is most of it. On the other named axis, Teasing, +0.257 became +0.248.
Quality. Mean predicted overall quality across the segments went 2.68 → 2.93 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.302 → 0.496+0.195identity cos neighbours 0.302 → 0.579d_b rescored +0.267 → +0.194d_a rescored +0.257 → +0.248d_a mined 0.257d_b mined 0.267min_cos_consec (site) 0.7429min_cos_anchor (site) 0.7429dataset emolialang enspeaker EN_B00014_S01943total 31.9schain gain +1.1 dBseam step 0.6 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, normally alert, average clarity, light breath
(slow, slightly relaxed, fairly steady, casual)It's a meme that references the event.
full caption & clip details
A young adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, no background noise; genuineness 2.9/6; vocal-burst blend 3.6/10; 3.3s, EN.
EN_B00014_S01943_W000112 · in -16.3 dBFS · gain -3.7 dB · emolia-00524
(sexual lust, hope enthusiasm optimism·normal-paced, neutral tension, moderately variable, casual)Well hold on, okay, so do you guys want to appeal this to a higher court? He's overlooking the information. I'm reviewing right now, hold on.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as sexual lust, hope enthusiasm optimism; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.2/6; vocal-burst blend 2.5/10; 8.1s, EN.
EN_B00014_S01943_W000113 · in -17.0 dBFS · gain -3.0 dB · emolia-00524
(amusement, teasing, intoxication altered states of consciousness· normal-paced, neutral tension, moderately variable, casual)It's actually not about the mystery, it's more about the Arby's. It's this (ahem) uh, Vince McMahon thing. It goes Ian x Australian girl, he's interested. Ian x mystery DM girl, he's orgasming. Ian x Arby's drive-thru guy, (chuckle) fully engaged. Yeaah.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as amusement, teasing, intoxication altered states of consciousness; style: casual, playful; average recording, quiet background; mildly explicit content; genuineness 4.4/6; vocal-burst blend 0.9/10; 20.9s, EN.
EN_B00014_S01943_W000114 · in -18.2 dBFS · gain -1.8 dB · emolia-00524
This chain comes from the two-sided rule: it only counts if both emotions move — Teasing down and Amusement up — by at least 0.25 each.
The chain starts with Amusement clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.29.
At the same time Teasing goes the other way, from 0.74 (higher than 74 % of clips in this corpus) to 0.99 (higher than 99 % of clips in this corpus), a change of +0.25. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.11, then +0.18, then +0.00 — a plateau around step 3, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.20 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.17 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.20, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 25 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.158 before conversion and 0.562 after — it rose by 0.404. Neighbour-to-neighbour the worst pair went 0.158 → 0.619. (The earlier render, with segment 1 left raw, scores 0.477 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Amusement moved +0.291 in the original and +0.264 after conversion — 91 % of the delta retained, which is essentially all of it. On the other named axis, Teasing, +0.251 became +0.208.
Quality. Mean predicted overall quality across the segments went 2.41 → 2.71 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.158 → 0.562+0.404identity cos neighbours 0.158 → 0.619d_b rescored +0.291 → +0.264d_a rescored +0.251 → +0.208d_a mined 0.251d_b mined 0.291min_cos_consec (site) 0.1718min_cos_anchor (site) 0.2005dataset emolialang enspeaker EN_AhcNfPVyYpytotal 24.3schain gain +1.8 dBseam step 2.1 dBcrossfades 100/100/100 ms
Script — 4 chunks, 4 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, balanced body, normal-paced, moderately variable, some disfluency
(normally alert, slightly relaxed, average clarity, casual)And eliminations (ahem) show my eliminations or team eliminations?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; no dominant emotion; style: casual, conversational; below-average recording, quiet background; genuineness 5.3/6; vocal-burst blend 4.3/10; 4.1s, EN.
EN_AhcNfPVyYpy_W000653 · in -18.7 dBFS · gain -1.3 dB · emolia-00687
(astonishment surprise, impatience and irritability· normally alert, neutral tension, average clarity, casual)(ahem) You got 17. You got 17 that game. No. Yeah.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as astonishment surprise, impatience and irritability; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.5/6; vocal-burst blend 0.0/10; 5.0s, EN.
EN_AhcNfPVyYpy_W000654 · in -19.4 dBFS · gain -0.6 dB · emolia-00687
(astonishment surprise, intoxication altered states of consciousness, embarrassment·energised, neutral tension, somewhat unclear, casual)When I started spectating you, you got 15, you got 2 kills since I started spectating you. Oh (low mumble) my. Uuuh. Interesting.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as astonishment surprise, intoxication altered states of consciousness, embarrassment; style: casual, playful; below-average recording, some background noise; mildly explicit content; genuineness 6.0/6; vocal-burst blend 0.6/10; 9.0s, EN.
EN_AhcNfPVyYpy_W000655 · in -19.3 dBFS · gain -0.7 dB · emolia-00687
(amusement, intoxication altered states of consciousness, teasing· energised, neutral tension, somewhat unclear, conversational)You best, you best be on tomorrow, bro. You best be on. (childlike giggle) I'll try, I'll try, I'll try. Yeah, yeah, yeah, sure, sure, sure.
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as amusement, intoxication altered states of consciousness, teasing; style: conversational, casual; below-average recording, some background noise; mildly explicit content; genuineness 5.9/6; vocal-burst blend 2.7/10; 6.5s, EN.
EN_AhcNfPVyYpy_W000661 · in -19.3 dBFS · gain -0.7 dB · emolia-00687
This chain comes from the two-sided rule: it only counts if both emotions move — Teasing down and Malevolence Malice up — by at least 0.25 each.
The chain starts with Malevolence Malice clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.37.
At the same time Teasing goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.91 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 44 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.820 before conversion and 0.825 after — it rose by 0.004. Neighbour-to-neighbour the worst pair went 0.898 → 0.853. (The earlier render, with segment 1 left raw, scores 0.701 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Malevolence Malice moved +0.366 in the original and +0.588 after conversion — 161 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Teasing, -0.259 became -0.613.
Quality. Mean predicted overall quality across the segments went 3.10 → 3.14 (+0.04) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.820 → 0.825+0.004identity cos neighbours 0.898 → 0.853d_b rescored +0.366 → +0.588d_a rescored -0.259 → -0.613d_a mined -0.259d_b mined 0.366min_cos_consec (site) 0.9100min_cos_anchor (site) 0.8842dataset emolialang enspeaker EN_MSVirmpAc3ctotal 43.1schain gain +2.0 dBseam step 0.7 dBcrossfades 100/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert, slightly relaxed
(teasing, intoxication altered states of consciousness, concentration · some disfluency, average clarity, light breath, casual)(low mumble) And there is a revoke authorization and revoke badge. If a badge is mistakenly awarded to someone, you can remove it. But if you remove it, there is a badge called Consolation Prize that you award them. You award them a badge as an apology for removing their badge.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as teasing, intoxication altered states of consciousness, concentration; style: casual, monologue; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 2.3/10; 19.7s, EN.
EN_MSVirmpAc3c_W000441 · in -18.4 dBFS · gain -1.6 dB · emolia-02129
(sourness, confusion·frequent disfluency, somewhat unclear, normal breath, casual)And you can revoke authorization if you accidentally award the wrong username or maybe say some group, the leadership of some team changes and they don't want
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, slightly guarded; reads as sourness, confusion; style: casual, monologue; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 1.7/10; 13.9s, EN.
EN_MSVirmpAc3c_W000442 · in -17.8 dBFS · gain -2.2 dB · emolia-02129
(malevolence malice, emotional numbness, contempt·some disfluency, average clarity, light breath, authoritative)Justin to be able to award the badge anymore, they want Miro to award the badge. You can remove the authorization from the old person and give it to the new person.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as malevolence malice, emotional numbness, contempt; style: authoritative, casual; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 1.3/10; 9.9s, EN.
EN_MSVirmpAc3c_W000443 · in -18.2 dBFS · gain -1.8 dB · emolia-02129