c-podcast-AB2 — voice-corrected

Corpus podcast in isolation, rule AB2.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_c-podcast-AB2.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
50segments re-voiced
0.701 → 0.805median worst-to-anchor identity cosine
79 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Confusion ↓  /  Contentmentidentity +0.19 emotion 165 %   c-podcast-AB2 · #1

This chain comes from the two-sided rule: it only counts if both emotions move — Confusion down and Contentment up — by at least 0.25 each.

The chain starts with Contentment clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.27.

At the same time Confusion goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.03, then +0.24 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.57 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.74 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.57, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 38 s · pt · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.572 before conversion and 0.760 after — it rose by 0.188. Neighbour-to-neighbour the worst pair went 0.701 → 0.790. (The earlier render, with segment 1 left raw, scores 0.617 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.256 in the original and +0.422 after conversion — 165 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Confusion, -0.275 became +0.029.

Quality. Mean predicted overall quality across the segments went 2.91 → 3.15 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.572 → 0.760 +0.188identity cos neighbours 0.701 → 0.790d_b rescored +0.256 → +0.422d_a rescored -0.275 → +0.029d_a mined -0.276d_b mined 0.270min_cos_consec (site) 0.7362min_cos_anchor (site) 0.5680dataset podcastlang ptspeaker 56246total 37.6schain gain +0.5 dBseam step 5.2 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-bright, fairly smooth, quiet background, normally alert, neutral tension, moderately variable, some disfluency, average clarity
(confusion, longing, pain · brisk, dramatic, playful) Pensando na cabeça, fazendo,
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as confusion, longing, pain; style: dramatic, playful; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 4.5/10; 5.1s, PT.
56246_00262716 · in -17.9 dBFS · gain -2.1 dB · podcast-02628
(shame, jealousy and envy, interest · brisk, conversational, casual) que vai encaixar aqui. Ah, tá, tem essa solução aqui. Ah, meu Deus, será que eu vou esquecer até chegar inglês? Cluck, cluc, clue, the time. 24 hours per year I was with my (low mumble) cabin. And this exige um espaço mental,
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, slightly guarded; reads as shame, jealousy and envy, interest; style: conversational, casual; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 5.6/10; 22.1s, PT.
56246_00263224 · in -18.1 dBFS · gain -1.9 dB · podcast-02284
(contentment, embarrassment, infatuation · normal-paced, casual, conversational) I pergunte the Gracia Infinita, which is this David Foster Wallace, which is a little more difficult to.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contentment, embarrassment, infatuation; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 7.1/10; 10.7s, PT.
56246_00269340 · in -16.2 dBFS · gain -3.8 dB · podcast-02630
Contempt ↓  /  Doubtidentity +0.60 emotion 77 %   c-podcast-AB2 · #2

This chain comes from the two-sided rule: it only counts if both emotions move — Contempt down and Doubt up — by at least 0.25 each.

The chain starts with Doubt clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.29.

At the same time Contempt goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are -0.10, then +0.22, then +0.17 — not a clean run: step 1 moves back the other way by 0.10 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.04 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.08 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.04, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 76 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.041 before conversion and 0.640 after — it rose by 0.600. Neighbour-to-neighbour the worst pair went 0.082 → 0.598. (The earlier render, with segment 1 left raw, scores 0.485 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.287 in the original and +0.220 after conversion — 77 % of the delta retained, which is most of it. On the other named axis, Contempt, -0.355 became -0.114.

Quality. Mean predicted overall quality across the segments went 2.99 → 3.31 (+0.31) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.041 → 0.640 +0.600identity cos neighbours 0.082 → 0.598d_b rescored +0.287 → +0.220d_a rescored -0.355 → -0.114d_a mined -0.356d_b mined 0.292min_cos_consec (site) 0.0789min_cos_anchor (site) 0.0418dataset podcastlang enspeaker 582211total 75.3schain gain +3.7 dBseam step 2.0 dBcrossfades 150/100/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, brisk, some disfluency, average clarity, light breath
(contempt, teasing, sexual lust · normally alert, slightly relaxed, moderately variable, casual) (low mumble) (low mumble) So yeah, yet again, the brilliant American voters, you know, who think of this as some kind of iterative game where they can just, you know, like punish the elites, punish the Democrats with their (low mumble) uh their protest votes, they're gonna get uh (ahem) they're gonna get a nice rude awakening. It's not gonna be fun.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contempt, teasing, sexual lust; style: casual, conversational; good recording, quiet background; genuineness 4.4/6; vocal-burst blend 7.4/10; 24.2s, EN.
582211_00172820 · in -27.9 dBFS · gain +7.9 dB · podcast-06225
(amusement, awe, interest · energised, neutral tension, moderately variable, casual) It's not gonna be nice. Yeah, I read an article in the New York Times a couple of days ago, and it was it was about this very thing where (low mumble) uh a large swath of Muslims, even in Dearborn, Michigan or Michigan or wherever that protest vote was going to be, went for Trump. And then when asked why they did that, it was the answer came back. Well, he said he wants peace. And I
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as amusement, awe, interest; style: casual, conversational; average recording, some background noise; genuineness 4.6/6; vocal-burst blend 8.3/10; 19.4s, EN.
582211_00175240 · in -24.7 dBFS · gain +4.7 dB · podcast-06224
(interest, amusement, astonishment surprise · energised, neutral tension, moderately variable, casual) was like, and and then but then they asked, they're they were like, (low mumble) um, well, why didn't you? I mean, Harris obviously wants speech. They're like, Well, she didn't have a plan though. I'm like, Trump didn't have a plan either. His whole plan was I want he actually said this verbatim, I want uh (ahem) Bibi Netanyahu to end this war quickly. What does okay, those words you can interpret that in many ways?
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as interest, amusement, astonishment surprise; style: casual, conversational; average recording, quiet background; genuineness 5.6/6; vocal-burst blend 10.0/10; 21.0s, EN.
582211_00177184 · in -25.0 dBFS · gain +5.0 dB · podcast-05554
(doubt, astonishment surprise, confusion · normally alert, slightly relaxed, fairly steady, casual) It's just it's it's just obviously that's not what he's calling for. But what what who knows what he's calling for? I mean, I think he would be that you wouldn't be getting like restrictions on which weapons the United States sends to Israel under a Trump administration. I promise you. I promise
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, astonishment surprise, confusion; style: casual, conversational; good recording, no background noise; genuineness 4.8/6; vocal-burst blend 6.2/10; 11.1s, EN.
582211_00179960 · in -29.1 dBFS · gain +9.1 dB · podcast-05543
Astonishment Surprise ↓  /  Emotional Numbnessidentity +0.11 emotion 83 %   c-podcast-AB2 · #3

This chain comes from the two-sided rule: it only counts if both emotions move — Astonishment Surprise down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, virtually no clip in this corpus scores higher. That is a total rise of 0.31.

At the same time Astonishment Surprise goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.60. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.09, then +0.18, then +0.04 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.77 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.76 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.77, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 44 s · fr · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.713 before conversion and 0.822 after — it rose by 0.109. Neighbour-to-neighbour the worst pair went 0.749 → 0.804. (The earlier render, with segment 1 left raw, scores 0.628 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.308 in the original and +0.257 after conversion — 83 % of the delta retained, which is most of it. On the other named axis, Astonishment Surprise, -0.600 became -0.604.

Quality. Mean predicted overall quality across the segments went 2.95 → 3.20 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.713 → 0.822 +0.109identity cos neighbours 0.749 → 0.804d_b rescored +0.308 → +0.257d_a rescored -0.600 → -0.604d_a mined -0.601d_b mined 0.310min_cos_consec (site) 0.7567min_cos_anchor (site) 0.7692dataset podcastlang frspeaker 930090total 42.8schain gain -0.1 dBseam step 4.7 dBcrossfades 150/100/150 ms
Script — 4 chunks, 4 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, balanced body, quiet background, fairly steady, frequent disfluency, somewhat unclear
(astonishment surprise, doubt, elation · normal-paced, very low-energy, neutral tension, casual) (ahem) The great of my life, I think. Okay, I'll recognize this story. I was here, I've had 10 years. I was in vaccine in Britain with my (ahem) parents, and we were in a location. At the epoch, it was more than
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as astonishment surprise, doubt, elation; style: casual, monologue; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 4.9/10; 16.2s, FR.
930090_00463596 · in -21.3 dBFS · gain +1.3 dB · podcast-04696
(slow, very low-energy, relaxed, conversational) (ahem) a (low mumble) prison of location, (low mumble) and in (low mumble) reality, they had (low mumble)
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, narrow pitch range, audible breath; affect is neutral, submissive, neutral openness; no dominant emotion; style: conversational, casual; below-average recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.0/10; 8.8s, FR.
930090_00465212 · in -21.7 dBFS · gain +1.7 dB · podcast-04458
(emotional numbness, longing · normal-paced, normally alert, neutral tension, conversational) not sure materialist at the moment. The only value materialist I have is my video that I'm off (low mumble) in the moon, you know. (low mumble) It's enormous. But
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, longing; style: conversational, casual; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 4.8/10; 10.4s, FR.
930090_00467243 · in -20.4 dBFS · gain +0.4 dB · podcast-04459
(emotional numbness, disgust, sourness · measured, normally alert, slightly relaxed, casual) video that has the value for me. The rest, (ahem) a bouteille, it has a lot of value. It's a boutique at 100 balls. Okay.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, disgust, sourness; style: casual, monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.9/10; 8.0s, FR.
930090_00468384 · in -21.9 dBFS · gain +1.9 dB · podcast-05087
Relief ↓  /  Interestidentity −0.07 emotion 95 %   c-podcast-AB2 · #4

This chain comes from the two-sided rule: it only counts if both emotions move — Relief down and Interest up — by at least 0.25 each.

The chain starts with Interest around average — 0.49, right about the corpus median — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.44.

At the same time Relief goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.21, then +0.12, then +0.10 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 80 s · es · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.878 before conversion and 0.807 after — it fell by 0.071. Neighbour-to-neighbour the worst pair went 0.878 → 0.807. (The earlier render, with segment 1 left raw, scores 0.753 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.439 in the original and +0.415 after conversion — 95 % of the delta retained, which is essentially all of it. On the other named axis, Relief, -0.268 became -0.225.

Quality. Mean predicted overall quality across the segments went 2.79 → 3.24 (+0.46) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.878 → 0.807 -0.071identity cos neighbours 0.878 → 0.807d_b rescored +0.439 → +0.415d_a rescored -0.268 → -0.225d_a mined -0.275d_b mined 0.439min_cos_consec (site) 0.9505min_cos_anchor (site) 0.9388dataset podcastlang esspeaker 1573total 78.8schain gain +0.0 dBseam step 1.9 dBcrossfades 100/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, balanced body, average recording, quiet background, normally alert, fairly steady, somewhat unclear, moderate pitch range
(relief, triumph, thankfulness gratitude · normal-paced, slightly relaxed, some disfluency, monologue) Y por eso, por eso falló aquel día. Claro, si tú no haces eso, porque no pones mecanismos de almacenamiento, no pones esas cosas, pues al final, por lo que tú decías antes, de la nevera y solo comes carne, pues al final, a la larga tienes problemas. Y bueno, yo creo que esto hay que (low mumble) replantear. Nos lo tenemos que replantear, no tenemos que tener muchas veces. hay cierta psicosis, cierta miedo con la nuclear. Creo que vamos a vivir de aquí a 2035
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as relief, triumph, thankfulness gratitude; style: monologue, casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 8.4/10; 29.6s, ES.
1573_00218592 · in -18.5 dBFS · gain -1.5 dB · podcast-00920
(fear, helplessness, hope enthusiasm optimism · measured, slightly relaxed, frequent disfluency, casual) si siguen los politicos que tenemos vamos a seguir. No los políticos me refiero a las personas, sino esa corriente. Vamos a seguir (low mumble) a sus
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, helplessness, hope enthusiasm optimism; style: casual, monologue; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 3.6/10; 10.2s, ES.
1573_00221548 · in -19.2 dBFS · gain -0.8 dB · podcast-00945
(hope enthusiasm optimism, triumph · normal-paced, neutral tension, some disfluency, monologue) demonization continua de la energía nuclear para crear esa psicosis anda sensación in the population that we have to start in contract of the nuclear.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism, triumph; style: monologue, casual; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 10.0/10; 17.7s, ES.
1573_00222560 · in -19.4 dBFS · gain -0.6 dB · podcast-00942
(interest, jealousy and envy · normal-paced, neutral tension, some disfluency, casual) In the mix tend to start laser, yo no estoy in contrables, we have to take the photovoltaica, the eólica, los de ciclo combinado, la de concentración, tenemos que tener todas las technologies and
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, jealousy and envy; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 10.0/10; 21.9s, ES.
1573_00224323 · in -19.7 dBFS · gain -0.3 dB · podcast-00936
Thankfulness Gratitude ↓  /  Jealousy and Envyidentity +0.32 emotion 80 %   c-podcast-AB2 · #5

This chain comes from the two-sided rule: it only counts if both emotions move — Thankfulness Gratitude down and Jealousy and Envy up — by at least 0.25 each.

The chain starts with Jealousy and Envy clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.28.

At the same time Thankfulness Gratitude goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.08 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.40 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.65 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.40, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 58 s · es · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.419 before conversion and 0.735 after — it rose by 0.316. Neighbour-to-neighbour the worst pair went 0.653 → 0.749. (The earlier render, with segment 1 left raw, scores 0.530 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Jealousy and Envy moved +0.280 in the original and +0.224 after conversion — 80 % of the delta retained, which is most of it. On the other named axis, Thankfulness Gratitude, -0.260 became -0.077.

Quality. Mean predicted overall quality across the segments went 2.79 → 3.17 (+0.38) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.419 → 0.735 +0.316identity cos neighbours 0.653 → 0.749d_b rescored +0.280 → +0.224d_a rescored -0.260 → -0.077d_a mined -0.260d_b mined 0.275min_cos_consec (site) 0.6480min_cos_anchor (site) 0.4019dataset podcastlang esspeaker 345026total 57.8schain gain +2.2 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, balanced body, quiet background, neutral tension, fairly steady, frequent disfluency, somewhat unclear, moderate pitch range
(thankfulness gratitude, contentment, affection · normal-paced, normally alert, casual, conversational) Muchas gracias, maestro, (low mumble) doctor Fragoso Cruz. Ando el uso de la voz (low mumble) para que justamente empecemos a introducirnos a esta función jurisdiccional a partir
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, contentment, affection; style: casual, conversational; below-average recording, quiet background; genuineness 5.8/6; vocal-burst blend 10.0/10; 13.1s, ES.
345026_00063056 · in -20.7 dBFS · gain +0.7 dB · podcast-04749
(triumph, shame, relief · normal-paced, normally alert, monologue, casual) término de qué es la justicia. (ahem) Magistrado Arturo, si usted nos pudiera apoyar ahora con su intervención, por favor. Claro que sí. Desde luego, muy buenas noches a los colegas andas que estamos dentro del ámbito jurisdiccional and the public who está escuchando. Tema muy importante, ¿no?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as triumph, shame, relief; style: monologue, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 8.0/10; 22.2s, ES.
345026_00064616 · in -21.3 dBFS · gain +1.3 dB · podcast-00012
(jealousy and envy, bitterness, anger · measured, subdued, casual, monologue) Porque hace poco nuestro presidente en México, para que lo sepa nuestro (ahem) colega in Colombia, decía que hay que los jueces, debemos aplicar justicia y no el derecho, ¿no? Andamos a esa dinámica de qué es la justicia, ¿no? (ahem) Y sin entrar a una definición tal, digamos, en primer punto, que es un derecho humano,
full caption & clip details
An adult masculine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, bitterness, anger; style: casual, monologue; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 7.5/10; 22.8s, ES.
345026_00066840 · in -21.9 dBFS · gain +1.9 dB · podcast-01959
Embarrassment ↓  /  Hope Enthusiasm Optimismidentity +0.38 emotion 71 %   c-podcast-AB2 · #6

This chain comes from the two-sided rule: it only counts if both emotions move — Embarrassment down and Hope Enthusiasm Optimism up — by at least 0.25 each.

The chain starts with Hope Enthusiasm Optimism clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.37.

At the same time Embarrassment goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.14, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.39 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.53 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.39, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 29 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.442 before conversion and 0.820 after — it rose by 0.378. Neighbour-to-neighbour the worst pair went 0.511 → 0.727. (The earlier render, with segment 1 left raw, scores 0.769 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.368 in the original and +0.263 after conversion — 71 % of the delta retained, which is most of it. On the other named axis, Embarrassment, -0.287 became -0.190.

Quality. Mean predicted overall quality across the segments went 2.80 → 2.98 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.442 → 0.820 +0.378identity cos neighbours 0.511 → 0.727d_b rescored +0.368 → +0.263d_a rescored -0.287 → -0.190d_a mined -0.282d_b mined 0.367min_cos_consec (site) 0.5303min_cos_anchor (site) 0.3895dataset podcastlang enspeaker 905162total 28.8schain gain +1.5 dBseam step 1.6 dBcrossfades 100/100 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, average recording, moderately variable
(embarrassment, amusement · slow, normally alert, relaxed, casual) And it also (low mumble) um it like they clearly are trying to show like the Chicago
full caption & clip details
A young adult feminine voice; delivery is normally alert, slow, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as embarrassment, amusement; style: casual, conversational; average recording, no background noise; genuineness 3.8/6; vocal-burst blend 2.1/10; 7.4s, EN.
905162_00040672 · in -25.9 dBFS · gain +5.9 dB · podcast-01825
(amusement, teasing, pleasure ecstasy · brisk, energised, neutral tension, casual) areas and stuff. And then I guess they like they're also having like (ahem) a this guy, I think one wrote for the Sun Times, one wrote for the tribunes. So it was like (ahem) a whose newspaper's gonna sell. Yeah, they're they're kinda like playing with the like a little bickering that
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as amusement, teasing, pleasure ecstasy; style: casual, playful; average recording, quiet background; mildly explicit content; genuineness 5.4/6; vocal-burst blend 8.0/10; 14.6s, EN.
905162_00041416 · in -22.2 dBFS · gain +2.2 dB · podcast-01853
(hope enthusiasm optimism, elation · normal-paced, normally alert, relaxed, casual) Uh (low mumble) but yeah, so we're gonna I think what we'll do is we're gonna listen to their review of Showgirls.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as hope enthusiasm optimism, elation; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 4.5/10; 7.0s, EN.
905162_00043152 · in -26.7 dBFS · gain +6.7 dB · podcast-01811
Fear ↓  /  Prideidentity −0.04 emotion REVERSED   c-podcast-AB2 · #7

This chain comes from the two-sided rule: it only counts if both emotions move — Fear down and Pride up — by at least 0.25 each.

The chain starts with Pride clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.34.

At the same time Fear goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.42 (lower than 58 % of clips in this corpus), a change of -0.44. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.12 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.73 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.73, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 27 s · it · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.756 before conversion and 0.720 after — it fell by 0.036. Neighbour-to-neighbour the worst pair went 0.756 → 0.720. (The earlier render, with segment 1 left raw, scores 0.699 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Pride moved +0.341 in the original and -0.025 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Fear, -0.443 became -0.103.

Quality. Mean predicted overall quality across the segments went 2.56 → 3.09 (+0.53) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.756 → 0.720 -0.036identity cos neighbours 0.756 → 0.720d_b rescored +0.341 → -0.025d_a rescored -0.443 → -0.103d_a mined -0.441d_b mined 0.338min_cos_consec (site) 0.8065min_cos_anchor (site) 0.7257dataset podcastlang itspeaker 549291total 26.5schain gain +1.8 dBseam step 2.1 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child masculine voice · neutral-toned, neutral-bright, fairly smooth, average recording, quiet background, normally alert, slightly relaxed, fairly steady
(normal-paced, no disfluency, wide pitch range, playful) vai in quei gruppi lì e tendi a fare il normale,
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, wide pitch range, light breath; affect is positive, neutral stance, neutral openness; no dominant emotion; style: playful, casual; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 2.6/10; 3.2s, IT.
549291_00038776 · in -21.8 dBFS · gain +1.8 dB · podcast-02228
(shame, embarrassment, interest · brisk, some disfluency, moderate pitch range, monologue) ti dicono tipo cose come oh, ma sei normale fra o. Ma in realtà sono loro che non sono normali. Si mettono una cazzo di maschera e fanno finta di essere le persone che non sono solo perché magari hanno incontrato un rapper, un figo che hanno ascoltato, e dicono: Minchia, quello lì è forte. Devo essere come lui. E allora tutti si sono ammassati. E vogliono essere tutti come quel determinato rapper,
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, embarrassment, interest; style: monologue, casual; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 10.0/10; 20.3s, IT.
549291_00039096 · in -22.3 dBFS · gain +2.3 dB · podcast-02222
(pride, infatuation, triumph · fast, some disfluency, wide pitch range, casual) determinato insomma, qualunque influencer che porta contenuti,
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as pride, infatuation, triumph; style: casual, playful; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 5.4/10; 3.4s, IT.
549291_00041120 · in -25.7 dBFS · gain +5.7 dB · podcast-02218
Hope Enthusiasm Optimism ↓  /  Shameidentity +0.03 emotion 89 %   c-podcast-AB2 · #8

This chain comes from the two-sided rule: it only counts if both emotions move — Hope Enthusiasm Optimism down and Shame up — by at least 0.25 each.

The chain starts with Shame around average — 0.55, higher than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.38.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.67 (higher than 67 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.75 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.75 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.75, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 52 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.689 before conversion and 0.716 after — it rose by 0.027. Neighbour-to-neighbour the worst pair went 0.734 → 0.758. (The earlier render, with segment 1 left raw, scores 0.628 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.410 in the original and +0.365 after conversion — 89 % of the delta retained, which is most of it. On the other named axis, Hope Enthusiasm Optimism, -0.324 became -0.237.

Quality. Mean predicted overall quality across the segments went 2.65 → 2.94 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.689 → 0.716 +0.027identity cos neighbours 0.734 → 0.758d_b rescored +0.410 → +0.365d_a rescored -0.324 → -0.237d_a mined -0.317d_b mined 0.380min_cos_consec (site) 0.7516min_cos_anchor (site) 0.7516dataset podcastlang enspeaker 664862total 51.0schain gain +4.5 dBseam step 0.8 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult somewhat masculine voice · slightly bright, neutral tension
(hope enthusiasm optimism, sexual lust, astonishment surprise · brisk, energised, moderately variable, casual) been to the shop, she's been to the gym, she's made breakfast, and she's even studied. That was just my brain telling me if I didn't do those things, I was going to
full caption & clip details
A young adult somewhat masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism, sexual lust, astonishment surprise; style: casual, playful; below-average recording, quiet background; mildly explicit content; genuineness 4.9/6; vocal-burst blend 10.0/10; 7.6s, EN.
664862_00016960 · in -22.5 dBFS · gain +2.5 dB · podcast-01909
(elation, astonishment surprise, amusement · brisk, energised, volatile, casual) fail at life. But you don't know, my heart was pounding, my armpits were itching from nerves. That was anxiety. Oh so mine, I have ENTP, oh, ENTP, which is a debater, is
full caption & clip details
A child feminine voice; delivery is energised, brisk, neutral tension, volatile; timbre is slightly cool, slightly bright, very rough, slightly thin; slurred, frequent disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as elation, astonishment surprise, amusement; style: casual, dramatic; below-average recording, some background noise; genuineness 4.2/6; vocal-burst blend 4.2/10; 17.5s, EN.
664862_00017712 · in -20.4 dBFS · gain +0.4 dB · podcast-00459
(shame, thankfulness gratitude, disgust · normal-paced, very low-energy, moderately variable, casual) (ahem) Um, what is it? Personality type, quick witted and audacious. People with the ENTP personality aren't afraid to disagree with the status quo. In fact, they're not afraid to disagree with pretty much anything or anyone. A few things light up these personalities more than a bit of verbal sparring. And if the conversation veers into a controversial terrain, so much the better. You think?
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is negative, slightly submissive, neutral openness; reads as shame, thankfulness gratitude, disgust; style: casual, dramatic; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 5.9/10; 26.3s, EN.
664862_00019680 · in -25.5 dBFS · gain +5.5 dB · podcast-06158
Anger ↓  /  Triumphidentity −0.17 emotion 90 %   c-podcast-AB2 · #9

This chain comes from the two-sided rule: it only counts if both emotions move — Anger down and Triumph up — by at least 0.25 each.

The chain starts with Triumph clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.39.

At the same time Anger goes the other way, from 0.88 (higher than 88 % of clips in this corpus) to 0.55 (higher than 55 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.18, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 29 s · ru · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.738 before conversion and 0.569 after — it fell by 0.169. Neighbour-to-neighbour the worst pair went 0.738 → 0.569. (The earlier render, with segment 1 left raw, scores 0.522 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Triumph moved +0.436 in the original and +0.394 after conversion — 90 % of the delta retained, which is essentially all of it. On the other named axis, Anger, -0.335 became -0.345.

Quality. Mean predicted overall quality across the segments went 2.81 → 3.08 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.738 → 0.569 -0.169identity cos neighbours 0.738 → 0.569d_b rescored +0.436 → +0.394d_a rescored -0.335 → -0.345d_a mined -0.336d_b mined 0.394min_cos_consec (site) 0.8830min_cos_anchor (site) 0.8361dataset podcastlang ruspeaker 286971total 28.6schain gain +2.3 dBseam step 0.4 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, average recording, quiet background, some disfluency, average clarity
(normal-paced, normally alert, slightly relaxed, casual) ебаная интеграция. Если Никита не в курсах,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 2.9/10; 3.0s, RU.
286971_00132264 · in -22.6 dBFS · gain +2.6 dB · podcast-00759
(interest, thankfulness gratitude, elation · brisk, normally alert, slightly relaxed, casual) были ребятки с углем. Не будут называть их названия. Но они очень хорошо они за рекламную интеграцию заплатили полтос и дали, соответственно, 100 килограмм угля. Я посидел в ванне, снял видосик, очень хорошо набрал. В итоге у них заказали угля.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as interest, thankfulness gratitude, elation; style: casual, dramatic; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 8.7/10; 14.6s, RU.
286971_00132568 · in -22.0 dBFS · gain +2.0 dB · podcast-00768
(triumph, pride, elation · fast, energised, neutral tension, casual) Опять же, другие чуваки, которые потом со мной связались, заказали угля на 7 миллионов рублей. И компания просто их прокидала, исчезла с рыбкой, забрав все бабки.
full caption & clip details
A young adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as triumph, pride, elation; style: casual, playful; average recording, quiet background; mildly explicit content; genuineness 5.7/6; vocal-burst blend 6.3/10; 11.3s, RU.
286971_00134024 · in -21.9 dBFS · gain +1.9 dB · podcast-00764
Emotional Numbness ↓  /  Contemplationidentity +0.04 emotion 209 %   c-podcast-AB2 · #10

This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Contemplation up — by at least 0.25 each.

The chain starts with Contemplation clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.25.

At the same time Emotional Numbness goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.63 (higher than 63 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.18, then +0.07 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 31 s · de · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.818 before conversion and 0.854 after — it rose by 0.036. Neighbour-to-neighbour the worst pair went 0.818 → 0.854. (The earlier render, with segment 1 left raw, scores 0.771 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.250 in the original and +0.523 after conversion — 209 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, -0.283 became -0.059.

Quality. Mean predicted overall quality across the segments went 2.84 → 3.12 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.818 → 0.854 +0.036identity cos neighbours 0.818 → 0.854d_b rescored +0.250 → +0.523d_a rescored -0.283 → -0.059d_a mined -0.283d_b mined 0.254min_cos_consec (site) 0.8655min_cos_anchor (site) 0.8299dataset podcastlang despeaker 225765total 30.6schain gain +0.2 dBseam step 2.5 dBcrossfades 100/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(emotional numbness · frequent disfluency, somewhat unclear, conversational, casual) Ja, ja, das (low mumble) muss ich sagen, das war doch dann auch (ahem) am weitesten weg meiner Weise. (ahem)
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as emotional numbness; style: conversational, casual; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 0.0/10; 6.3s, DE.
225765_00109440 · in -23.1 dBFS · gain +3.1 dB · podcast-05769
(helplessness, distress, sadness · some disfluency, somewhat unclear, casual, conversational) (ahem) Auch für mich privat und ich habe viel draus gelernt, aber (low mumble) muss sagen, also Rom ist schon. Wenn du dann aus dem Flugzeug steigst, in den Bus steigst und dann über die Autostrade auf Italienisch dann Autobahn fährst und so nach über einer Stunde dann ins Stadtzentrum ankommst,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness, distress, sadness; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 4.1/10; 16.6s, DE.
225765_00110072 · in -24.5 dBFS · gain +4.5 dB · podcast-03424
(contemplation, sadness, disappointment · some disfluency, average clarity, casual, conversational) (ahem) muss man schon sagen, der Römer damals hat schon viel geschaffen, was heute noch erhalten ist. Und das gibt es in Deutschland so nicht zu finden.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, sadness, disappointment; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 2.6/10; 8.1s, DE.
225765_00111728 · in -25.0 dBFS · gain +5.0 dB · podcast-05763
Pride ↓  /  Jealousy and Envyidentity +0.03 emotion 69 %   c-podcast-AB2 · #11

This chain comes from the two-sided rule: it only counts if both emotions move — Pride down and Jealousy and Envy up — by at least 0.25 each.

The chain starts with Jealousy and Envy clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.26.

At the same time Pride goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.63 (higher than 63 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.09 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 64 s · da · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.920 before conversion and 0.951 after — it rose by 0.031. Neighbour-to-neighbour the worst pair went 0.940 → 0.951. (The earlier render, with segment 1 left raw, scores 0.871 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Jealousy and Envy moved +0.265 in the original and +0.182 after conversion — 69 % of the delta retained. On the other named axis, Pride, -0.338 became -0.635.

Quality. Mean predicted overall quality across the segments went 3.18 → 3.44 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.920 → 0.951 +0.031identity cos neighbours 0.940 → 0.951d_b rescored +0.265 → +0.182d_a rescored -0.338 → -0.635d_a mined -0.344d_b mined 0.260min_cos_consec (site) 0.9401min_cos_anchor (site) 0.9205dataset podcastlang daspeaker 537315total 63.1schain gain +0.8 dBseam step 1.5 dBcrossfades 100/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, subdued, slightly relaxed, fairly steady
(pride, thankfulness gratitude, disappointment · measured, monologue) det jo, jeg synes også spændende, hvor vi jo snakke om, (low mumble) at det har et stort potentiel i skolen, at man jo lækker for mange underviser inspireret til, når de kan komme tilbage i skolen, at det kan være en måde at arbejde på også for børn der, og at man sagten skulle stille podcast opgaver.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, thankfulness gratitude, disappointment; style: monologue; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 0.6/10; 18.6s, DA.
537315_00116240 · in -26.5 dBFS · gain +6.5 dB · podcast-01886
(contentment, doubt, relief · normal-paced, monologue) Jeg synes, det var enormt spændende både at komme lidt (low mumble) med bord om (low mumble) linjerne ved corona for børn, for jeg synes, det er en podcast, jeg har opdaget den her tid, som jeg virkelig har nyt også og lyft til par episoder (low mumble) samt to dræge af dem og jo han. (low mumble) Og vi ikke få nogle gode samtaler på baggrund af det. Og jeg synes i hvert fald, at der er nog interessante ting der med og som familie både at.
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contentment, doubt, relief; style: monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 3.7/10; 22.0s, DA.
537315_00122511 · in -22.1 dBFS · gain +2.1 dB · podcast-01888
(jealousy and envy, longing, contentment · measured, monologue, whispered) starte et fælles projekt op i den her tid, (low mumble) og især når de nu er en familie, der har tiden til det. Og så er det der dilemma, der nogle gange kan være i forhold til, eller samspil med den læring, der så sker fra skolen af, hvor de her er i hvert fald en familie, der mixer lidt læringen, man lærer ved det konkrete projekt, og så den læring, der også er resultateret fra deres skoler af.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, longing, contentment; style: monologue, whispered; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 3.4/10; 22.9s, DA.
537315_00124708 · in -26.0 dBFS · gain +6.0 dB · podcast-01167
Emotional Numbness ↓  /  Affectionidentity +0.08 emotion REVERSED   c-podcast-AB2 · #12

This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Affection up — by at least 0.25 each.

The chain starts with Affection clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.30.

At the same time Emotional Numbness goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.57 (higher than 57 % of clips in this corpus), a change of -0.37. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.02, then +0.10, then -0.04 — not a clean run: step 4 moves back the other way by 0.04 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.71 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.73 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.71, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 37 s · da · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.751 before conversion and 0.830 after — it rose by 0.079. Neighbour-to-neighbour the worst pair went 0.765 → 0.741. (The earlier render, with segment 1 left raw, scores 0.697 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Affection moved +0.305 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Emotional Numbness, -0.400 became +0.032.

Quality. Mean predicted overall quality across the segments went 2.35 → 3.11 (+0.76) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.751 → 0.830 +0.079identity cos neighbours 0.765 → 0.741d_b rescored +0.305 → +0.000d_a rescored -0.400 → +0.032d_a mined -0.368d_b mined 0.303min_cos_consec (site) 0.7342min_cos_anchor (site) 0.7122dataset podcastlang daspeaker 476458total 36.1schain gain +0.8 dBseam step 2.3 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · slightly warm, average recording, very low-energy, relaxed, fairly narrow pitch
(emotional numbness · measured, moderately variable, frequent disfluency, whispered) Historien om hans døde på skiner har plet ham så meget. At denet bliver den baser.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is slightly warm, dark, smooth, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is positive, slightly submissive, neutral openness; reads as emotional numbness; style: whispered, ASMR; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 2.3/10; 6.1s, DA.
476458_00023496 · in -17.2 dBFS · gain -2.8 dB · podcast-00304
(contemplation, longing · measured, fairly steady, frequent disfluency, whispered) Hendes forældre var festlige mennesker. Der var interesseret i kunst og kultur.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is slightly warm, slightly dark, fairly smooth, slightly thin; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is mildly positive, submissive, neutral openness; reads as contemplation, longing; style: whispered, ASMR; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 2.8/10; 5.6s, DA.
476458_00042952 · in -18.8 dBFS · gain -1.2 dB · podcast-05558
(emotional numbness, longing · measured, fairly steady, frequent disfluency, whispered) Måske afbrød Grædent her.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is slightly warm, dark, slightly rough, thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, submissive, neutral openness; reads as emotional numbness, longing; style: whispered, storytelling; average recording, no background noise; genuineness 3.1/6; vocal-burst blend 4.9/10; 3.9s, DA.
476458_00044904 · in -18.2 dBFS · gain -1.8 dB · podcast-00291
(longing, contemplation, affection · slow, fairly steady, frequent disfluency, ASMR) Jeg mens slå de dateren værd alene hjemme. Og så opte hun sin måsts familie.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly positive, submissive, neutral openness; reads as longing, contemplation, affection; style: ASMR, whispered; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 1.7/10; 6.4s, DA.
476458_00045296 · in -17.7 dBFS · gain -2.3 dB · podcast-00295
(measured, fairly steady, little disfluency, whispered) Jeg vil gerne teget på træt af en pige. Men hendes voldsomme død overskyder for det levende mennesker. Når jøren fortæller jeg, at Græde kunne i, så han også, at hun fik så med sig i graven.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is slightly warm, slightly dark, slightly rough, slightly thin; slurred, little disfluency, fairly narrow pitch, audible breath; affect is mildly positive, submissive, neutral openness; no dominant emotion; style: whispered, ASMR; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 1.9/10; 14.8s, DA.
476458_00058376 · in -17.6 dBFS · gain -2.4 dB · podcast-00293
Pain ↓  /  Confusionidentity −0.05 emotion 77 %   c-podcast-AB2 · #13

This chain comes from the two-sided rule: it only counts if both emotions move — Pain down and Confusion up — by at least 0.25 each.

The chain starts with Confusion clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.27.

At the same time Pain goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.07, then +0.20 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 50 s · da · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.858 before conversion and 0.803 after — it fell by 0.055. Neighbour-to-neighbour the worst pair went 0.858 → 0.803. (The earlier render, with segment 1 left raw, scores 0.633 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.268 in the original and +0.207 after conversion — 77 % of the delta retained, which is most of it. On the other named axis, Pain, -0.400 became -0.383.

Quality. Mean predicted overall quality across the segments went 3.27 → 3.39 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.858 → 0.803 -0.055identity cos neighbours 0.858 → 0.803d_b rescored +0.268 → +0.207d_a rescored -0.400 → -0.383d_a mined -0.400d_b mined 0.269min_cos_consec (site) 0.9500min_cos_anchor (site) 0.9482dataset podcastlang daspeaker 322938total 49.5schain gain +0.6 dBseam step 1.1 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, slightly relaxed, fairly steady, somewhat unclear
(pain, intoxication altered states of consciousness, jealousy and envy · normal-paced, subdued, some disfluency, monologue) Her i huset er det sådan, at (low mumble) vi holder ikke af at købe vores marmelade nogen steder. Vi laver den selv. Vi har alle (ahem) frugter og bærer og hvad vi skal. Det har vi i haver, vi har det levende hegn rundt omkring. Jeg har selv plantet omkring 10.000 planter. Og mange af dem blomstår, det er også godt forbi det. (low mumble) Men noget af det giver brumbær, noget af det giver svæskblom, og noget af det giver nogle andre blommer. Nu giver Kirsbær, og nu giver ting, som vi tager hen. Så laver vi et (low mumble) parmel tit og lidt og dat. Og i stedet for sukker, så bruger vi et
full caption & clip details
A young adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as pain, intoxication altered states of consciousness, jealousy and envy; style: monologue, casual; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 8.1/10; 29.7s, DA.
322938_00089480 · in -24.4 dBFS · gain +4.4 dB · podcast-06376
(disgust · normal-paced, normally alert, some disfluency, casual) det synes vi jo smert vildt godt. Og så kigger vi på den og sige, vi spiser lidt, for det her er sømer. Og hånd skal man ikke bruge så meget som sukker, så derfor kan vi roligt lige drødst
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 9.2/10; 9.7s, DA.
322938_00092592 · in -24.7 dBFS · gain +4.7 dB · podcast-04693
(confusion · measured, normally alert, frequent disfluency, monologue) Sådan lever vi godt, og igen. Vi vil have noget ordentligt. Vi skal ikke have særlig meget. Vi er ikke pensionister, men vi vil bare have et ordentligt liv. Hvis
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as confusion; style: monologue, whispered; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 3.3/10; 10.4s, DA.
322938_00093616 · in -25.0 dBFS · gain +5.0 dB · podcast-04718
Jealousy and Envy ↓  /  Shameidentity +0.56 emotion REVERSED   c-podcast-AB2 · #14

This chain comes from the two-sided rule: it only counts if both emotions move — Jealousy and Envy down and Shame up — by at least 0.25 each.

The chain starts with Shame clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.27.

At the same time Jealousy and Envy goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are -0.02, then +0.23, then +0.06 — not a clean run: step 1 moves back the other way by 0.02 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.33 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.26 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.33, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 85 s · pt · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.296 before conversion and 0.857 after — it rose by 0.561. Neighbour-to-neighbour the worst pair went 0.318 → 0.858. (The earlier render, with segment 1 left raw, scores 0.740 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

The emotional move did not survive. Re-scored end to end, Shame moved +0.274 in the original and -0.558 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Jealousy and Envy, -0.337 became -0.475.

Quality. Mean predicted overall quality across the segments went 2.89 → 3.44 (+0.55) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.296 → 0.857 +0.561identity cos neighbours 0.318 → 0.858d_b rescored +0.274 → -0.558d_a rescored -0.337 → -0.475d_a mined -0.344d_b mined 0.273min_cos_consec (site) 0.2602min_cos_anchor (site) 0.3256dataset podcastlang ptspeaker 416864total 84.5schain gain +3.0 dBseam step 1.3 dBcrossfades 100/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · quiet background, fairly steady
(jealousy and envy, interest, contemplation · brisk, normally alert, slightly relaxed, monologue) (surprised gasp) Ah, Splendor isso, né? Tem as magias curativas. Então, ao mesmo tempo que você tem essa combinação de dois, que é um pouco limitante pras classes, porque elas vão se repetindo, você acaba tendo muito mais liberdade com uma classe diferente do que você estava planejando fazer as coisas que você gostaria que o seu personagem fizesse. O que é outra coisa que enriquece muito a narrativa, porque você não vai estar engessado do tipo, ah, se eu vou fazer um bárbaro, meu bárbaro tem que ser aquele
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as jealousy and envy, interest, contemplation; style: monologue, authoritative; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 7.4/10; 28.6s, PT.
416864_00226719 · in -17.1 dBFS · gain -2.9 dB · podcast-02082
(doubt, jealousy and envy, contemplation · normal-paced, normally alert, slightly relaxed, monologue) cara bruto, aquele cara burro, aquele cara que dá porrada. Ah, mas dá pra fazer de outra forma. Dá, mas é forçar um pouco a barra do sistema também. Aqui, você tem essa liberdade de poder não só ter as habilidades que fogem do padrão de não ser só algo de combate, mas você conseguir, putz, eu quero muito o aspecto do bardo de ser o maluco do combate ali, que vai ter as habilidades com arma, vai ser um pouquinho mais tanker, mas eu queria
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, jealousy and envy, contemplation; style: monologue, authoritative; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 6.0/10; 27.0s, PT.
416864_00229573 · in -17.7 dBFS · gain -2.3 dB · podcast-02078
(interest, concentration, pride · normal-paced, normally alert, slightly relaxed, casual) que ele fizesse essa outra coisa. Então por que você não pega a classe que tem esse domínio do combate, mas ela também tem esse outro domínio aqui que faz umas acrobacias, umas coisas mais diferentes. E pra mim, é uma das coisas mais geniais do Daggerheart, é essa brincadeira que ele faz com os domínios e as classes.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, concentration, pride; style: casual, conversational; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 6.0/10; 16.1s, PT.
416864_00232270 · in -18.3 dBFS · gain -1.7 dB · podcast-02097
(shame, contemplation, helplessness · measured, very low-energy, relaxed, casual) Acho que adicionando nisso que você tava falando do bárbaro, né? E de misturar as classes, novamente a coisa de cartas que só servem pra narrativa, né? Tem no. Eu acho que é no Blade. Não, qual que é o vermelho mesmo?
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is warm, dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as shame, contemplation, helplessness; style: casual, storytelling; below-average recording, quiet background; genuineness 3.5/6; vocal-burst blend 4.9/10; 13.3s, PT.
416864_00233904 · in -21.6 dBFS · gain +1.6 dB · podcast-02559
Confusion ↓  /  Contentmentidentity +0.08 emotion 113 %   c-podcast-AB2 · #15

This chain comes from the two-sided rule: it only counts if both emotions move — Confusion down and Contentment up — by at least 0.25 each.

The chain starts with Contentment clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.29.

At the same time Confusion goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.62 (higher than 62 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.11, then +0.18 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.78 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.82 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.78, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 26 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.648 before conversion and 0.732 after — it rose by 0.084. Neighbour-to-neighbour the worst pair went 0.648 → 0.732. (The earlier render, with segment 1 left raw, scores 0.607 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.289 in the original and +0.326 after conversion — 113 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Confusion, -0.276 became -0.212.

Quality. Mean predicted overall quality across the segments went 2.74 → 2.87 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.648 → 0.732 +0.084identity cos neighbours 0.648 → 0.732d_b rescored +0.289 → +0.326d_a rescored -0.276 → -0.212d_a mined -0.277d_b mined 0.289min_cos_consec (site) 0.8155min_cos_anchor (site) 0.7784dataset podcastlang enspeaker 149941total 25.2schain gain +5.1 dBseam step 1.1 dBcrossfades 100/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, average clarity, light breath
(moderately variable, some disfluency, wide pitch range, casual) Before everything is a C Merg plot. (ahem) Uh yeah. Okay, so
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, conversational; average recording, no background noise; genuineness 5.6/6; vocal-burst blend 3.0/10; 3.1s, EN.
149941_00616528 · in -22.8 dBFS · gain +2.8 dB · podcast-03916
(embarrassment · fairly steady, frequent disfluency, moderate pitch range, casual) on (ahem) the hold. Yes. So uh (low mumble) then well well, after the last thing we talked about, but before the C Mergullaby, Taylor goes and visits all those other people that were guiding figures for her.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as embarrassment; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 4.5/10; 12.0s, EN.
149941_00616928 · in -22.1 dBFS · gain +2.1 dB · podcast-03917
(contentment, affection · moderately variable, frequent disfluency, wide pitch range, casual) Yes. (ahem) Um so Charlotte and the kids and (ahem) um and Forrest. Uh (ahem) and then Glenn and Kaye, and then finally (low mumble) um writing that note to Miss Militia.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contentment, affection; style: casual, playful; good recording, quiet background; genuineness 4.1/6; vocal-burst blend 4.7/10; 10.4s, EN.
149941_00618136 · in -23.4 dBFS · gain +3.4 dB · podcast-03929
Contentment ↓  /  Contemplationidentity +0.45 emotion 106 %   c-podcast-AB2 · #16

This chain comes from the two-sided rule: it only counts if both emotions move — Contentment down and Contemplation up — by at least 0.25 each.

The chain starts with Contemplation clearly present — 0.69, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.30.

At the same time Contentment goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.67 (higher than 67 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.14, then +0.16 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.35 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.08 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.35, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 73 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.435 before conversion and 0.890 after — it rose by 0.454. Neighbour-to-neighbour the worst pair went 0.100 → 0.836. (The earlier render, with segment 1 left raw, scores 0.682 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.301 in the original and +0.319 after conversion — 106 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contentment, -0.300 became -0.285.

Quality. Mean predicted overall quality across the segments went 3.12 → 3.26 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.435 → 0.890 +0.454identity cos neighbours 0.100 → 0.836d_b rescored +0.301 → +0.319d_a rescored -0.300 → -0.285d_a mined -0.305d_b mined 0.303min_cos_consec (site) 0.0781min_cos_anchor (site) 0.3502dataset podcastlang enspeaker 548382total 72.0schain gain +2.8 dBseam step 2.9 dBcrossfades 150/100 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, quiet background, normal-paced, average clarity
(contentment, fatigue exhaustion, thankfulness gratitude · normally alert, neutral tension, moderately variable, casual) It is. Even marathon training it (ahem) for for me was a lot. So you bring up a good point.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contentment, fatigue exhaustion, thankfulness gratitude; style: casual, playful; good recording, quiet background; genuineness 4.4/6; vocal-burst blend 6.0/10; 20.1s, EN.
548382_00327892 · in -30.5 dBFS · gain +10.5 dB · podcast-06063
(fear, relief, infatuation · very low-energy, slightly relaxed, fairly steady, casual) (low mumble) Um for someone who, you know, is constantly working through their anxiety. How is this time in your head when you're out doing these longer runs? Or you know, I guess when you're swimming, you're so focused on, you know, your arms, your legs, and your breathing. So there's not much else to think about in between. But otherwise, when you are exercising, how's the how how are you in your head?
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as fear, relief, infatuation; style: casual, conversational; good recording, quiet background; genuineness 2.7/6; vocal-burst blend 5.8/10; 25.4s, EN.
548382_00329904 · in -29.3 dBFS · gain +9.3 dB · podcast-00988
(contemplation, infatuation, awe · normally alert, neutral tension, moderately variable, casual) (ahem) Um, you know, it's, you know, as I've gotten more into meditation, it's kind of like the ongoing body scan where I really try to focus on literally what my muscles are doing at that time. (low mumble) Um, and it's not easy. I mean, like anybody who's tried meditating, like it is the it is the human brain's like tendency to just all of a sudden wonder uh (low mumble) how to spell cornucopia or
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contemplation, infatuation, awe; style: casual, playful; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 7.7/10; 26.8s, EN.
548382_00332480 · in -26.7 dBFS · gain +6.7 dB · podcast-06006
Hope Enthusiasm Optimism ↓  /  Longingidentity −0.03 emotion REVERSED   c-podcast-AB2 · #17

This chain comes from the two-sided rule: it only counts if both emotions move — Hope Enthusiasm Optimism down and Longing up — by at least 0.25 each.

The chain starts with Longing around average — 0.48, lower than 52 % of clips in this corpus — and ends with it clearly present at 0.73, higher than 73 % of clips in this corpus. That is a total rise of 0.26.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.57 (higher than 57 % of clips in this corpus), a change of -0.42. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.09, then +0.07, then +0.10 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 41 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.858 before conversion and 0.829 after — it fell by 0.029. Neighbour-to-neighbour the worst pair went 0.760 → 0.713. (The earlier render, with segment 1 left raw, scores 0.733 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The emotional move did not survive. Re-scored end to end, Longing moved +0.265 in the original and -0.250 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Hope Enthusiasm Optimism, -0.424 became -0.486.

Quality. Mean predicted overall quality across the segments went 3.00 → 3.07 (+0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.858 → 0.829 -0.029identity cos neighbours 0.760 → 0.713d_b rescored +0.265 → -0.250d_a rescored -0.424 → -0.486d_a mined -0.419d_b mined 0.256min_cos_consec (site) 0.8616min_cos_anchor (site) 0.8226dataset podcastlang enspeaker 501941total 39.8schain gain +3.6 dBseam step 1.1 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · fairly smooth, quiet background, normal-paced, normally alert, fairly steady, average clarity, light breath
(hope enthusiasm optimism, thankfulness gratitude, elation · slightly relaxed, some disfluency, moderate pitch range, conversational) aside from studying is we have a really good team. So my three partners, Dr. Kumbh has a start, my co-founder,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as hope enthusiasm optimism, thankfulness gratitude, elation; style: conversational, casual; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 3.1/10; 7.1s, EN.
501941_00103735 · in -18.2 dBFS · gain -1.8 dB · podcast-03404
(slightly relaxed, some disfluency, wide pitch range, casual) she's a biochemist uh by training. She has a postdoc from Stanford. Her previous business, she built an ELISA machine,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 2.0/10; 8.8s, EN.
501941_00104440 · in -17.7 dBFS · gain -2.3 dB · podcast-03403
(intoxication altered states of consciousness, embarrassment, jealousy and envy · neutral tension, some disfluency, moderate pitch range, casual) small machine in her her garage in New Zealand, and eventually build a company with her husband in is based in New Zealand. The company is over 20 years old (ahem) uh by now. So that was the first thing that she did. Then my other uh (ahem) partner, he's (low mumble) uh former athlete,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as intoxication altered states of consciousness, embarrassment, jealousy and envy; style: casual, monologue; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.5/10; 19.4s, EN.
501941_00105320 · in -17.7 dBFS · gain -2.3 dB · podcast-05137
(slightly relaxed, frequent disfluency, moderate pitch range, casual) for uh a (ahem) couple of decades, and finally our CTO who's just joined.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 3.0/10; 5.1s, EN.
501941_00107720 · in -16.5 dBFS · gain -3.5 dB · podcast-03400
Interest ↓  /  Impatience and Irritabilityidentity +0.46 emotion 81 %   c-podcast-AB2 · #18

This chain comes from the two-sided rule: it only counts if both emotions move — Interest down and Impatience and Irritability up — by at least 0.25 each.

The chain starts with Impatience and Irritability around average — 0.45, lower than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.54.

At the same time Interest goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.51 (higher than 51 % of clips in this corpus), a change of -0.49. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.19, then +0.11 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.24 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.28 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.24, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 92 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.297 before conversion and 0.760 after — it rose by 0.463. Neighbour-to-neighbour the worst pair went 0.281 → 0.671. (The earlier render, with segment 1 left raw, scores 0.517 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.546 in the original and +0.441 after conversion — 81 % of the delta retained, which is most of it. On the other named axis, Interest, -0.490 became -0.405.

Quality. Mean predicted overall quality across the segments went 2.99 → 3.21 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.297 → 0.760 +0.463identity cos neighbours 0.281 → 0.671d_b rescored +0.546 → +0.441d_a rescored -0.490 → -0.405d_a mined -0.489d_b mined 0.541min_cos_consec (site) 0.2796min_cos_anchor (site) 0.2401dataset podcastlang enspeaker 487739total 90.7schain gain +4.6 dBseam step 1.8 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, light breath
(interest, hope enthusiasm optimism, awe · normal-paced, normally alert, neutral tension, casual) but it it gives you a little bit of devil's advocate as I can see how Penske kind of fell into this trap they set for themselves, if that makes sense. So I feel like I took a little bit longer than needed to cover this, but at the same time, this is something that like took wave storms on Twitter. I'm I I mean like F1 people are commenting on this. This is one of the biggest cheating scandals we've ever seen for the biggest race in the world, and it's really complicated to go through multiple layers.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as interest, hope enthusiasm optimism, awe; style: casual, conversational; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 9.0/10; 26.1s, EN.
487739_00186066 · in -24.1 dBFS · gain +4.2 dB · podcast-02529
(relief, pain, pleasure ecstasy · brisk, energised, neutral tension, casual) So for those of you that are like, okay, you guys need to get start to start previewing the race. We I think we had to cover this in good detail before we start getting into a pretty comical guide that I made. So (low mumble) um Carson, from your point of view, do you have any final thoughts on this, or did that kind of paint a clear picture as to what happened? (low mumble)
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as relief, pain, pleasure ecstasy; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 9.4/10; 23.6s, EN.
487739_00188672 · in -21.7 dBFS · gain +1.7 dB · podcast-02525
(doubt, fear, confusion · measured, subdued, relaxed, casual) Well, in (low mumble) this (low mumble) particular instance,
full caption & clip details
An adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as doubt, fear, confusion; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 10.0/10; 28.0s, EN.
487739_00193060 · in -22.5 dBFS · gain +2.5 dB · podcast-02522
(impatience and irritability, anger, bitterness · normal-paced, normally alert, slightly relaxed, casual) Roger Penske was not, and same with the last one. Roger Penske was not involved in any of the punish the punishment giving. He was just he he put himself on the sidelines and told Doug, not told him, but Doug Bowles did not have any influence
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, anger, bitterness; style: casual; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 3.0/10; 13.6s, EN.
487739_00195855 · in -23.6 dBFS · gain +3.6 dB · podcast-02532
Concentration ↓  /  Affectionidentity +0.03 emotion REVERSED   c-podcast-AB2 · #19

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Affection up — by at least 0.25 each.

The chain starts with Affection clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.34.

At the same time Concentration goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.16, then +0.05, then +0.06, then +0.07 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 108 s · pt · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.803 before conversion and 0.832 after — it rose by 0.029. Neighbour-to-neighbour the worst pair went 0.817 → 0.744. (The earlier render, with segment 1 left raw, scores 0.706 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The emotional move did not survive. Re-scored end to end, Affection moved +0.337 in the original and -0.818 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Concentration, -0.334 became -0.169.

Quality. Mean predicted overall quality across the segments went 2.80 → 3.30 (+0.50) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.803 → 0.832 +0.029identity cos neighbours 0.817 → 0.744d_b rescored +0.337 → -0.818d_a rescored -0.334 → -0.169d_a mined -0.336d_b mined 0.342min_cos_consec (site) 0.8919min_cos_anchor (site) 0.8179dataset podcastlang ptspeaker 462285total 106.3schain gain +2.5 dBseam step 0.6 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a child feminine voice · neutral-bright, fairly smooth, average recording, quiet background, normal-paced
(concentration, contemplation · normally alert, neutral tension, moderately variable, casual) recursos tecnológicos. Prestou atenção neste detalhe? Desde o ensino fundamental, você tem acesso aos diversos gêneros textuais, sem ao menos perceber a argumentação, muitas vezes implícita em cada um deles. Como vimos no início desta unidade, (low mumble)
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration, contemplation; style: casual, playful; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 3.6/10; 16.4s, PT.
462285_00529248 · in -31.9 dBFS · gain +11.9 dB · podcast-02760
(concentration, relief, jealousy and envy · normally alert, slightly relaxed, fairly steady, casual) (ahem) em cada gênero pode coexistir várias tipologias textuais, sendo a dissertação argumentativa uma das mais recorrentes, visto que a maioria dos sujeitos tem a prática de externar seu ponto de vista e seus argumentos. (low mumble) Externar seu ponto de vista e seus argumentos. Como o advento das redes sociais é de praxe, percebemos comentários com alto teor e argumentativo, embora uma grande porcentagem desses internautas seja anônima.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, relief, jealousy and envy; style: casual, monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 5.1/10; 26.6s, PT.
462285_00530888 · in -31.9 dBFS · gain +11.9 dB · podcast-02804
(jealousy and envy, interest, fatigue exhaustion · normally alert, slightly relaxed, fairly steady, monologue) Os gêneros de humor, em geral, são argumentativos, pois envolvem questões políticas e militâncias diversas. Por isso, os cartoons da atualidade vêm sendo ressignificados por sujeitos mais reais, como é o caso dos memes. Para tanto os subgêneros dos quadrinhos como o cartão podem ser argumentativos por trás de uma narrativa curta, o que pode torná-los tão complexos quanto um texto mais longo, como artigo de opinião. (low mumble)
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as jealousy and envy, interest, fatigue exhaustion; style: monologue, narration; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 3.9/10; 29.8s, PT.
462285_00533544 · in -31.7 dBFS · gain +11.7 dB · podcast-02760
(triumph, shame, contentment · very low-energy, slightly relaxed, fairly steady, casual) Cartoon most de um senhor dialogando com um menino a respeito do primeiro dia de aula. (low mumble) A imagem é de um cartoon, o qual mostra dois personagens negros. Não tem nada de negro aqui. Um adulto e uma criança em uma sala de estar tendo as seguintes falas. Você pode me dizer detalhes sobre o seu primeiro dia na escola?
full caption & clip details
A young adult somewhat feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, frequent disfluency, fairly narrow pitch, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as triumph, shame, contentment; style: casual, whispered; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 3.5/10; 23.8s, PT.
462285_00543320 · in -33.5 dBFS · gain +13.5 dB · podcast-02757
(affection, contentment · normally alert, neutral tension, moderately variable, casual) Perguntou o senhor o menino, por sua vez responde. Não é uma grande novidade. Eu terei que retornar para lá amanhã novamente. O cartoon evidencia uma realidade que retrata o profundo.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as affection, contentment; style: casual, monologue; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 5.5/10; 10.5s, PT.
462285_00545696 · in -31.8 dBFS · gain +11.8 dB · podcast-02785
Astonishment Surprise ↓  /  Disgustidentity −0.15 emotion REVERSED   c-podcast-AB2 · #20

This chain comes from the two-sided rule: it only counts if both emotions move — Astonishment Surprise down and Disgust up — by at least 0.25 each.

The chain starts with Disgust clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.28.

At the same time Astonishment Surprise goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.03 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.59 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.71 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.59, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 35 s · es · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.473 before conversion and 0.318 after — it fell by 0.155. Neighbour-to-neighbour the worst pair went 0.473 → 0.318. (The earlier render, with segment 1 left raw, scores 0.396 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

The emotional move did not survive. Re-scored end to end, Disgust moved +0.277 in the original and -0.487 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Astonishment Surprise, -0.271 became -0.475.

Quality. Mean predicted overall quality across the segments went 2.64 → 2.88 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.473 → 0.318 -0.155identity cos neighbours 0.473 → 0.318d_b rescored +0.277 → -0.487d_a rescored -0.271 → -0.475d_a mined -0.271d_b mined 0.276min_cos_consec (site) 0.7135min_cos_anchor (site) 0.5928dataset podcastlang esspeaker 492833total 34.4schain gain +4.3 dBseam step 1.9 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · neutral-toned, fairly smooth, moderately variable
(astonishment surprise, doubt, confusion · slow, normally alert, slightly relaxed, conversational) Es que te juro que tiene.
full caption & clip details
A child feminine voice; delivery is normally alert, slow, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; crisply articulate, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly submissive, neutral openness; reads as astonishment surprise, doubt, confusion; style: conversational, casual; average recording, no background noise; genuineness 2.4/6; vocal-burst blend 2.9/10; 3.3s, ES.
492833_00497135 · in -29.1 dBFS · gain +9.1 dB · podcast-01989
(disappointment, jealousy and envy, sourness · normal-paced, energised, neutral tension, casual) (low mumble) La sensibilidad de una película de Navidad de Netflix, ¿vale? O sea, y es que es muy mala. O sea, es muy mala. Y. Y no sé, o sea, y el final de la película es como, bueno, pues ya está, no he modificado nada y nada cambia. Y entonces, el Batman que aparece, en vez de ser Ben Affleck, es. (low mumble) Ay, George Clooney.
full caption & clip details
A child feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as disappointment, jealousy and envy, sourness; style: casual, dramatic; below-average recording, quiet background; genuineness 4.8/6; vocal-burst blend 6.8/10; 26.0s, ES.
492833_00497460 · in -23.2 dBFS · gain +3.2 dB · podcast-04032
(disgust, amusement, infatuation · fast, normally alert, slightly relaxed, conversational) eso es lo. la gracieta, ¿no? Y a él se le cae el diente que se había pegado con Super Gloop. (contented sigh) O
full caption & clip details
A young adult feminine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as disgust, amusement, infatuation; style: conversational, casual; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 5.0/10; 5.4s, ES.
492833_00500248 · in -25.3 dBFS · gain +5.3 dB · podcast-04767