This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_c-podcast-PXR.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words
are a real non-speech sound, printed where it happens.
The full generated caption for any clip is under “full caption & clip
details”. Its perceived-gender and background-noise clauses were re-rendered from
the numeric buckets, because the versions stored in the corpus index had those two
ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
The tags and captions describe the ORIGINAL clips — they are the
corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the
converted audio are the ones printed in each card's own paragraph and chips.
This chain comes from the proxy rule: the same two-sided test as above, but because Confusion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Confusion clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.34.
At the same time Longing goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.13 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.28 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.35 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.28, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 17 s · es · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.227 before conversion and 0.748 after — it rose by 0.522. Neighbour-to-neighbour the worst pair went 0.291 → 0.748. (The earlier render, with segment 1 left raw, scores 0.551 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.339 in the original and +0.226 after conversion — 67 % of the delta retained. On the other named axis, Longing, -0.271 became -0.075.
Quality. Mean predicted overall quality across the segments went 2.41 → 2.92 (+0.51) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.227 → 0.748+0.522identity cos neighbours 0.291 → 0.748d_b rescored +0.339 → +0.226d_a rescored -0.271 → -0.075d_a mined -0.267d_b mined 0.340min_cos_consec (site) 0.3478min_cos_anchor (site) 0.2767dataset podcastlang esspeaker 610445total 16.2schain gain +0.9 dBseam step 0.9 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, normally alert, moderately variable, some disfluency, wide pitch range, light breath
(longing, shame · fast, slightly relaxed, average clarity, casual)de los tiempos que están cambiando, ¿cachai? Y todavía se producen estas contradicciones media extrañas.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is neutral-toned, dark, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as longing, shame; style: casual, playful; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 8.0/10; 4.0s, ES.
610445_00182568 · in -24.4 dBFS · gain +4.4 dB · podcast-04063
(teasing, jealousy and envy, amusement· fast, slightly relaxed, average clarity, casual)Es que la mayoría de la gente que yo conozco que consume gentai son mujeres.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as teasing, jealousy and envy, amusement; style: casual, playful; average recording, quiet background; genuineness 5.5/6; vocal-burst blend 6.7/10; 3.8s, ES.
610445_00183048 · in -22.6 dBFS · gain +2.6 dB · podcast-04068
(confusion, thankfulness gratitude, distress·normal-paced, neutral tension, somewhat unclear, casual)De una sí que llega una mina con su esposo como una cuestión. Y la mina como que va a conocer al casero. Y el casero ahí la agarra, güey. La duerma, así que la
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as confusion, thankfulness gratitude, distress; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 6.0/6; vocal-burst blend 9.1/10; 8.8s, ES.
610445_00186536 · in -23.5 dBFS · gain +3.5 dB · podcast-00653
This chain comes from the proxy rule: the same two-sided test as above, but because Shame is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Shame clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.26.
At the same time Longing goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.19, then +0.01, then +0.07 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.76 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.76 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.76, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 46 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.545 before conversion and 0.488 after — it fell by 0.057. Neighbour-to-neighbour the worst pair went 0.650 → 0.535. (The earlier render, with segment 1 left raw, scores 0.387 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.264 in the original and +0.253 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Longing, -0.302 became -0.254.
Quality. Mean predicted overall quality across the segments went 2.98 → 3.04 (+0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.545 → 0.488-0.057identity cos neighbours 0.650 → 0.535d_b rescored +0.264 → +0.253d_a rescored -0.302 → -0.254d_a mined -0.299d_b mined 0.265min_cos_consec (site) 0.7633min_cos_anchor (site) 0.7633dataset podcastlang enspeaker 531228total 44.6schain gain +2.8 dBseam step 1.1 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, fairly steady
(longing, contemplation, pain · normal-paced, normally alert, slightly relaxed, casual)life with. Remember we talked about the traditional wedding vows
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as longing, contemplation, pain; style: casual, monologue; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 3.3/10; 3.2s, EN.
531228_00171024 · in -28.6 dBFS · gain +8.6 dB · podcast-00871
(infatuation, sexual lust, longing ·slow, normally alert, relaxed, casual)it begins by I take you to be my wedded, you know, wife,
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as infatuation, sexual lust, longing; style: casual, conversational; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 6.2/10; 4.7s, EN.
531228_00171520 · in -29.9 dBFS · gain +9.9 dB · podcast-00887
(infatuation, contemplation, pain·measured, subdued, slightly relaxed, monologue)husband. The the whole point of that one is remember that that that's all about me acknowledging that this is my free choice, that I haven't been forced into doing this, that I don't believe that, oh, this is the magical one. Because why that's so important is at the end of the day when Aaron and I go through the inevitable challenges and struggles and the good, the bad, the ugly, all of it, that every single moment I want to go,
full caption & clip details
An infant masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, minimal breath; affect is mildly negative, neutral stance, slightly guarded; reads as infatuation, contemplation, pain; style: monologue, narration; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 3.9/10; 27.2s, EN.
531228_00171984 · in -29.9 dBFS · gain +9.9 dB · podcast-06080
(shame, sexual lust, malevolence malice· measured, very low-energy, relaxed, monologue)I picked you, I chose you, and therefore I'm responsible also for the state of this marriage. And and (ahem) I this is not just going, well,
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, neutral openness; reads as shame, sexual lust, malevolence malice; style: monologue, casual; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 5.4/10; 10.1s, EN.
531228_00174696 · in -30.7 dBFS · gain +10.7 dB · podcast-05213
This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Interest around average — 0.57, higher than 57 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.43.
At the same time Anger goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.59 (higher than 59 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.09, then +0.10, then +0.24 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.20 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.20 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.20, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 91 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.154 before conversion and 0.611 after — it rose by 0.458. Neighbour-to-neighbour the worst pair went 0.165 → 0.646. (The earlier render, with segment 1 left raw, scores 0.449 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.435 in the original and +0.523 after conversion — 120 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Anger, -0.309 became -0.146.
Quality. Mean predicted overall quality across the segments went 2.93 → 3.22 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.154 → 0.611+0.458identity cos neighbours 0.165 → 0.646d_b rescored +0.435 → +0.523d_a rescored -0.309 → -0.146d_a mined -0.314d_b mined 0.430min_cos_consec (site) 0.2034min_cos_anchor (site) 0.2021dataset podcastlang enspeaker 407303total 89.9schain gain +2.4 dBseam step 1.6 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-bright, normal-paced, some disfluency
(anger · normally alert, neutral tension, moderately variable, casual)Not much person, it consumed 1400 botellas of this medication.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as anger; style: casual, conversational; below-average recording, some background noise; genuineness 5.7/6; vocal-burst blend 3.7/10; 7.8s, EN.
407303_00120728 · in -13.1 dBFS · gain -6.9 dB · podcast-02934
(jealousy and envy, bitterness, disappointment·subdued, neutral tension, fairly steady, monologue)the mandibula is bastard.
full caption & clip details
An adolescent masculine voice; delivery is subdued, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as jealousy and envy, bitterness, disappointment; style: monologue, casual; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 8.3/10; 29.8s, EN.
407303_00123064 · in -18.1 dBFS · gain -1.9 dB · podcast-06464
(disappointment, triumph, sadness· subdued, slightly relaxed, fairly steady, monologue)Al final era una soberana (low mumble) estupidez. O sea, de vez en cuando sigue pasando. No penséis que esto sucedió hace un siglo y ya (low mumble) se acabó. No que va de vez en cuando van a seguir sacando productos milagrosos, y yo creo que el único producto milagroso que hay en la vida, sobre todo en la salud, es cuidarse,
full caption & clip details
A young adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, triumph, sadness; style: monologue, casual; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 9.5/10; 22.9s, EN.
407303_00145359 · in -17.1 dBFS · gain -2.9 dB · podcast-06469
(interest, hope enthusiasm optimism, affection· subdued, slightly relaxed, fairly steady, monologue)tener buenos hábitos, moverse, hacer ejercicio, y esto va a hacer que estemos mucho más sanos. Cualquier (ahem) panacea que nos vendan siempre desde mi punto de vista (low mumble) va a ser falsa. Pero bueno, una auténtica, una auténtica locura (low mumble) dentro de la medicina que duró 10 años. O sea, imaginaos la de miles de personas que tuvieron repercusiones (ahem) con esto. Vamos a pasar a la siguiente historia, y esto es una historia de la que yo hablé hace muchos años con amigos míos psicólogos,
full caption & clip details
An adolescent masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, hope enthusiasm optimism, affection; style: monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 8.7/10; 29.9s, EN.
407303_00147664 · in -17.8 dBFS · gain -2.2 dB · podcast-06461
This chain comes from the proxy rule: the same two-sided test as above, but because Doubt is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Doubt clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.29.
At the same time Contemplation goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.09, then +0.20 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.25 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.17 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.25, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 46 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.221 before conversion and 0.574 after — it rose by 0.353. Neighbour-to-neighbour the worst pair went 0.082 → 0.548. (The earlier render, with segment 1 left raw, scores 0.617 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.289 in the original and +0.154 after conversion — 53 % of the delta retained. On the other named axis, Contemplation, -0.254 became -0.249.
Quality. Mean predicted overall quality across the segments went 2.64 → 3.05 (+0.41) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.221 → 0.574+0.353identity cos neighbours 0.082 → 0.548d_b rescored +0.289 → +0.154d_a rescored -0.254 → -0.249d_a mined -0.252d_b mined 0.294min_cos_consec (site) 0.1672min_cos_anchor (site) 0.2526dataset podcastlang enspeaker 422727total 45.7schain gain +4.0 dBseam step 2.0 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an elderly masculine voice · quiet background
(contemplation, affection, intoxication altered states of consciousness · measured, very low-energy, relaxed, casual)The more the more you think about it, it's just like parents are like that. Because parents will walk in there, or even like any adult just to kill time. And you know, it's my I'm obligated to like greet people, you know, I like enjoy doing that. I enjoy starting conversations with strangers,
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, relaxed, steady; timbre is slightly cool, dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, submissive, slightly guarded; reads as contemplation, affection, intoxication altered states of consciousness; style: casual, whispered; below-average recording, quiet background; genuineness 4.4/6; vocal-burst blend 4.1/10; 17.3s, EN.
422727_00180528 · in -50.3 dBFS · gain +30.3 dB · podcast-05808
(disappointment, sourness, contemplation · measured, subdued, relaxed, whispered)But you know, when I'm sitting there and it's like, oh I don't want help from you, or they just don't say anything to me, but they'll only speak kindly to the white people that work there. Which is great, and like they just never say anything to me. I'm like, that's cool. Like, you know. I've had a few instances where I've been called out for being my skin tone or having my hair. And it's great.
full caption & clip details
A young adult somewhat masculine voice; delivery is subdued, measured, relaxed, steady; timbre is slightly cool, dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, slightly guarded; reads as disappointment, sourness, contemplation; style: whispered, monologue; below-average recording, quiet background; genuineness 4.0/6; vocal-burst blend 5.4/10; 24.6s, EN.
422727_00182456 · in -49.1 dBFS · gain +29.1 dB · podcast-05794
(doubt, embarrassment, helplessness·normal-paced, normally alert, neutral tension, casual)I feel like someone was being racist towards me on the job, I don't know what I would do. I really
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as doubt, embarrassment, helplessness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 6.0/6; vocal-burst blend 5.6/10; 4.1s, EN.
422727_00186584 · in -35.1 dBFS · gain +15.1 dB · podcast-05799
This chain comes from the proxy rule: the same two-sided test as above, but because Elation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Elation clearly present — 0.69, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.25.
At the same time Astonishment Surprise goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.01 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.19 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.16 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.19, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 48 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.174 before conversion and 0.677 after — it rose by 0.503. Neighbour-to-neighbour the worst pair went 0.154 → 0.812. (The earlier render, with segment 1 left raw, scores 0.559 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Elation moved +0.251 in the original and +0.060 after conversion — 24 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Astonishment Surprise, -0.299 became -0.723.
Quality. Mean predicted overall quality across the segments went 2.72 → 3.15 (+0.43) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.174 → 0.677+0.503identity cos neighbours 0.154 → 0.812d_b rescored +0.251 → +0.060d_a rescored -0.299 → -0.723d_a mined -0.299d_b mined 0.251min_cos_consec (site) 0.1613min_cos_anchor (site) 0.1862dataset podcastlang enspeaker 798328total 47.2schain gain +4.6 dBseam step 2.1 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-bright, slightly rough, balanced body, normal-paced, some disfluency, wide pitch range
(astonishment surprise, triumph, jealousy and envy · very low-energy, slightly relaxed, moderately variable, casual)Yeah, we did. We did see a John Starks jersey at the CNE after the TFD. No, it was before. It was just you and me. In games four and five, though, this is what's crazy. Starks had double digit fourth quarters. And in game six, he had a sixteen point fourth quarter. And he was all pretty much dare I say dominant. (low mumble) Um in the finals, and then again, kind of like the fall from Grace like Tiger, the sh the shit from Grace.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as astonishment surprise, triumph, jealousy and envy; style: casual, conversational; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 3.5/10; 25.6s, EN.
798328_00339160 · in -29.2 dBFS · gain +9.2 dB · podcast-01530
(amusement, disappointment, jealousy and envy · very low-energy, neutral tension, moderately variable, casual)because he like props to him. He just wouldn't stop. He just kept trying. Because he's got it. He's a shooter. Like he's got to help his team. And it was just sad. It's like, oh, some like, please, please, it's already dead. John Starks.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as amusement, disappointment, jealousy and envy; style: casual, playful; average recording, quiet background; mildly explicit content; genuineness 6.0/6; vocal-burst blend 6.9/10; 14.4s, EN.
798328_00342216 · in -27.6 dBFS · gain +7.6 dB · podcast-05705
(elation·energised, fully relaxed, fairly steady, casual)talking about the AFC championship Cleveland Browns versus Denver Broncos in nineteen eighty five and nineteen eighty six. Oh yeah. Consecutive years too.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, fully relaxed, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as elation; style: casual, playful; poor recording, some background noise; explicit content; genuineness 3.4/6; vocal-burst blend 2.2/10; 7.6s, EN.
798328_00344335 · in -24.0 dBFS · gain +4.0 dB · podcast-05709
This chain comes from the proxy rule: the same two-sided test as above, but because Hope Enthusiasm Optimism is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Hope Enthusiasm Optimism clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.31.
At the same time Doubt goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.55 (higher than 55 % of clips in this corpus), a change of -0.43. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.09, then +0.22 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.62 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.56 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.62, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 67 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.609 before conversion and 0.805 after — it rose by 0.196. Neighbour-to-neighbour the worst pair went 0.560 → 0.805. (The earlier render, with segment 1 left raw, scores 0.538 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.312 in the original and +0.352 after conversion — 113 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Doubt, -0.427 became -0.360.
Quality. Mean predicted overall quality across the segments went 3.01 → 3.24 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.609 → 0.805+0.196identity cos neighbours 0.560 → 0.805d_b rescored +0.312 → +0.352d_a rescored -0.427 → -0.360d_a mined -0.427d_b mined 0.314min_cos_consec (site) 0.5564min_cos_anchor (site) 0.6173dataset podcastlang enspeaker 467152total 66.6schain gain +4.4 dBseam step 1.0 dBcrossfades 150/100 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert, neutral tension, average clarity
(doubt, interest · brisk, fairly steady, some disfluency, casual)like Jefferson shut down or else he's not gonna perform. So I get it for this year, but like you said, it's like is this defense good enough to go into a Super Bowl and win, or even to just go through that NFC (low mumble) uh in the playoffs? Probably not. This defense isn't really there yet. But right now, you know, they like the way their offense is playing and they went for it.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as doubt, interest; style: casual, playful; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 9.1/10; 21.3s, EN.
467152_00028132 · in -24.8 dBFS · gain +4.8 dB · podcast-06333
(amusement, intoxication altered states of consciousness, infatuation·normal-paced, moderately variable, frequent disfluency, casual)So I thought it was a lot, but (ahem) I get it. (exhausted groan) Uh I'll talk real quick. You mentioned William Jackson III. Your boys, the Steelers picked up a good one on the cheap. We'll see if he's still (ahem) good.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as amusement, intoxication altered states of consciousness, infatuation; style: casual, conversational; good recording, quiet background; genuineness 5.4/6; vocal-burst blend 5.6/10; 20.2s, EN.
467152_00030256 · in -25.7 dBFS · gain +5.7 dB · podcast-06318
(hope enthusiasm optimism, elation, contentment· normal-paced, moderately variable, some disfluency, casual)Yeah, clearly a move for the future. Steelers aren't going anywhere this year, but I like what they're doing. They're starting to put the pieces in place. Uh, (exhausted groan) speaking of the Steelers, though, they weren't done. Chase Claypool, you heard the rumors who's going to get traded, and they found a trade partner in the Chicago Bears who gave up a 2023 second round pick for Chase Claypool. Steelers fan, how do you feel about it? And just for the Bears, what do you
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, elation, contentment; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 3.9/6; vocal-burst blend 5.7/10; 25.4s, EN.
467152_00035527 · in -24.6 dBFS · gain +4.6 dB · podcast-06311
This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Interest clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.30.
At the same time Contentment goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.07 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 65 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.873 before conversion and 0.781 after — it fell by 0.091. Neighbour-to-neighbour the worst pair went 0.873 → 0.748. (The earlier render, with segment 1 left raw, scores 0.615 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.301 in the original and +0.370 after conversion — 123 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contentment, -0.341 became -0.348.
Quality. Mean predicted overall quality across the segments went 2.58 → 3.27 (+0.69) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.873 → 0.781-0.091identity cos neighbours 0.873 → 0.748d_b rescored +0.301 → +0.370d_a rescored -0.341 → -0.348d_a mined -0.343d_b mined 0.297min_cos_consec (site) 0.8669min_cos_anchor (site) 0.8687dataset podcastlang enspeaker 462815total 64.4schain gain +4.9 dBseam step 0.5 dBcrossfades 150/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an elderly somewhat feminine voice · slightly cool, slightly dark, slightly thin, below-average recording, quiet background, very low-energy, frequent disfluency, somewhat unclear
(contentment, jealousy and envy, intoxication altered states of consciousness · measured, neutral tension, moderately variable, casual)Oh dear. But (low mumble) um, I had some family things going on. That's the reason why we (low mumble) um haven't got back to (low mumble) um our regularly scheduled program of the random horror show. And the last what (low mumble) um we talked about, just reiterating with you guys, is that (ahem) uh we were talking about Watchmen, (ahem) uh, which is the HBO one season wonder.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is slightly cool, slightly dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, slightly submissive, neutral openness; reads as contentment, jealousy and envy, intoxication altered states of consciousness; style: casual, ASMR; below-average recording, quiet background; genuineness 4.5/6; vocal-burst blend 5.7/10; 26.8s, EN.
462815_00017996 · in -28.2 dBFS · gain +8.2 dB · podcast-01870
(astonishment surprise, elation, pride·slow, relaxed, moderately variable, casual)(low mumble) Um I read that they're not going to be making (low mumble) um a second season of Watchmen. Watchmen is very, very good. I mean, I mean, it is so I mean, you know what? I'm just gonna say it's exactly
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is slightly cool, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, slightly submissive, neutral openness; reads as astonishment surprise, elation, pride; style: casual, conversational; below-average recording, quiet background; genuineness 4.9/6; vocal-burst blend 3.9/10; 16.7s, EN.
462815_00020672 · in -28.3 dBFS · gain +8.3 dB · podcast-01872
(interest, intoxication altered states of consciousness, pride ·measured, relaxed, fairly steady, casual)what's going on in our society and world, and I know Watchmen is science fiction, you know, (low mumble) uh kind of like thrown in (low mumble) um drama (low mumble) uh type of (ahem) um show because it's set in an alternate reality of ours, and (low mumble) um
full caption & clip details
An elderly feminine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is slightly cool, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, neutral openness; reads as interest, intoxication altered states of consciousness, pride; style: casual, monologue; below-average recording, quiet background; genuineness 4.5/6; vocal-burst blend 4.5/10; 21.3s, EN.
462815_00022336 · in -27.9 dBFS · gain +7.9 dB · podcast-01865
This chain comes from the proxy rule: the same two-sided test as above, but because Amusement is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Amusement below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.62.
At the same time Shame goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.62 (higher than 62 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.39, then +0.23 — a fairly even climb, though some clips carry more of the change than others.
The largest step is 0.39, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.
Same speaker? The least similar clip scores 0.77 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.83 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.77, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 36 s · es · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.765 before conversion and 0.516 after — it fell by 0.248. Neighbour-to-neighbour the worst pair went 0.773 → 0.516. (The earlier render, with segment 1 left raw, scores 0.534 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Amusement moved +0.622 in the original and +0.188 after conversion — 30 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Shame, -0.329 became -0.501.
Quality. Mean predicted overall quality across the segments went 2.78 → 3.02 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.765 → 0.516-0.248identity cos neighbours 0.773 → 0.516d_b rescored +0.622 → +0.188d_a rescored -0.329 → -0.501d_a mined -0.327d_b mined 0.622min_cos_consec (site) 0.8266min_cos_anchor (site) 0.7724dataset podcastlang esspeaker 45286total 35.2schain gain -0.7 dBseam step 4.3 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, quiet background, normally alert, some disfluency, light breath
(shame, sadness, distress · normal-paced, slightly relaxed, fairly steady, monologue)maestro de vida, ¿no? No tanto maestro of the color, but maestro of you. And it was the ultimate,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, sadness, distress; style: monologue; good recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.0/10; 6.6s, ES.
45286_00092700 · in -28.5 dBFS · gain +8.5 dB · podcast-02582
(affection, infatuation, triumph· normal-paced, slightly relaxed, fairly steady, monologue)And ellos, as you platic, you can have a little bit of camera record and manage a little bit more the situation. And if no, pues always enseñar. But like,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as affection, infatuation, triumph; style: monologue, casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 2.7/10; 16.2s, ES.
45286_00096644 · in -29.8 dBFS · gain +9.8 dB · podcast-02578
(amusement, astonishment surprise, confusion·brisk, neutral tension, moderately variable, casual)no solo team enseñan materias, concepts, formulas, but I think they enseñan cosas of the way. Okay, not divorciarte, (wistful sigh) right? But
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as amusement, astonishment surprise, confusion; style: casual, conversational; good recording, quiet background; genuineness 4.2/6; vocal-burst blend 3.9/10; 12.8s, ES.
45286_00098268 · in -26.6 dBFS · gain +6.6 dB · podcast-02552
This chain comes from the proxy rule: the same two-sided test as above, but because Impatience and Irritability is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Impatience and Irritability around average — 0.52, higher than 52 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.43.
At the same time Malevolence Malice goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.63 (higher than 63 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.22 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.78 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.70 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.78, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 28 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.669 before conversion and 0.654 after — it fell by 0.014. Neighbour-to-neighbour the worst pair went 0.635 → 0.700. (The earlier render, with segment 1 left raw, scores 0.599 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.434 in the original and +0.370 after conversion — 85 % of the delta retained, which is most of it. On the other named axis, Malevolence Malice, -0.370 became -0.840.
Quality. Mean predicted overall quality across the segments went 2.68 → 2.94 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.669 → 0.654-0.014identity cos neighbours 0.635 → 0.700d_b rescored +0.434 → +0.370d_a rescored -0.370 → -0.840d_a mined -0.357d_b mined 0.430min_cos_consec (site) 0.7026min_cos_anchor (site) 0.7777dataset podcastlang enspeaker 137155total 27.8schain gain +2.4 dBseam step 2.0 dBcrossfades 100/100 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · wide pitch range
(malevolence malice, contempt, anger · slow, very low-energy, slightly relaxed, casual)Yeah, yeah. So that actor, his wife was a congresswoman in the 1940s. (low mumble) Uh she actually is the one who ran against Richard Nixon in the 1950 Senate race. Uh
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as malevolence malice, contempt, anger; style: casual, didactic; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 0.3/10; 16.6s, EN.
137155_00363712 · in -22.8 dBFS · gain +2.8 dB · podcast-01558
(astonishment surprise, amusement, confusion·normal-paced, energised, neutral tension, casual)she was a big time Democrat and (low mumble) uh big a second. Is she is she not the one that there's rumors she may have had an affair
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as astonishment surprise, amusement, confusion; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 5.2/6; vocal-burst blend 2.3/10; 7.1s, EN.
137155_00365416 · in -15.4 dBFS · gain -4.6 dB · podcast-01563
(impatience and irritability, contempt· normal-paced, normally alert, neutral tension, casual)with LBJ? No, no, there's a rumors, you can find the letters. She had a
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as impatience and irritability, contempt; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.6/6; vocal-burst blend 4.2/10; 4.4s, EN.
137155_00366120 · in -21.0 dBFS · gain +1.0 dB · podcast-01557
This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Contemplation clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.32.
At the same time Hope Enthusiasm Optimism goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.10 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 52 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.755 before conversion and 0.795 after — it rose by 0.040. Neighbour-to-neighbour the worst pair went 0.755 → 0.795. (The earlier render, with segment 1 left raw, scores 0.709 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.314 in the original and +0.353 after conversion — 112 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Hope Enthusiasm Optimism, -0.293 became -0.294.
Quality. Mean predicted overall quality across the segments went 3.01 → 3.12 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.755 → 0.795+0.040identity cos neighbours 0.755 → 0.795d_b rescored +0.314 → +0.353d_a rescored -0.293 → -0.294d_a mined -0.294d_b mined 0.316min_cos_consec (site) 0.8277min_cos_anchor (site) 0.8337dataset podcastlang enspeaker 102629total 51.3schain gain +3.5 dBseam step 1.3 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult somewhat feminine voice · neutral-toned, fairly smooth, balanced body, some disfluency, average clarity, moderate pitch range, light breath
(hope enthusiasm optimism, concentration, pain · measured, subdued, slightly relaxed, whispered)And we can point to social neuroscience to help us understand this. Our brains are wired in this organizing principle of minimizing threat and maximizing reward. And this goes towards social behavior. So anytime we're social in groups, we want to minimize the threat and maximize the reward.
full caption & clip details
A young adult somewhat feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as hope enthusiasm optimism, concentration, pain; style: whispered, monologue; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 0.3/10; 24.0s, EN.
102629_00028027 · in -19.9 dBFS · gain -0.1 dB · podcast-06067
(concentration ·normal-paced, normally alert, neutral tension, casual)It is really designed to help people stay alive by quickly and easily remembering what is good and bad in the environment. And so
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as concentration; style: casual, conversational; good recording, no background noise; genuineness 3.6/6; vocal-burst blend 3.9/10; 9.0s, EN.
102629_00030424 · in -19.9 dBFS · gain -0.1 dB · podcast-05084
(contemplation, fear, helplessness· normal-paced, normally alert, slightly relaxed, casual)when you think about feedback, it feels painful, it feels threatening, and it feels like something we really want to avoid. So a lot of us avoid giving it. And it's also one of those things, most feedback is going to be constructive.
full caption & clip details
An adult somewhat feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contemplation, fear, helplessness; style: casual, monologue; good recording, quiet background; genuineness 3.1/6; vocal-burst blend 4.4/10; 18.7s, EN.
102629_00031320 · in -20.4 dBFS · gain +0.4 dB · podcast-05059
This chain comes from the proxy rule: the same two-sided test as above, but because Doubt is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Doubt clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.26.
At the same time Sexual Lust goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.63 (higher than 63 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.10 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.80 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.80 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.80, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 53 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.671 before conversion and 0.581 after — it fell by 0.091. Neighbour-to-neighbour the worst pair went 0.722 → 0.701. (The earlier render, with segment 1 left raw, scores 0.552 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
The emotional move did not survive. Re-scored end to end, Doubt moved +0.266 in the original and -0.065 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Sexual Lust, -0.328 became -0.503.
Quality. Mean predicted overall quality across the segments went 2.44 → 3.00 (+0.56) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.671 → 0.581-0.091identity cos neighbours 0.722 → 0.701d_b rescored +0.266 → -0.065d_a rescored -0.328 → -0.503d_a mined -0.332d_b mined 0.262min_cos_consec (site) 0.7958min_cos_anchor (site) 0.7958dataset podcastlang enspeaker 941719total 52.3schain gain +5.1 dBseam step 2.3 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · quiet background, moderately variable
(sexual lust · normal-paced, normally alert, relaxed, storytelling)Do you have goals? Like, are you comfortable with you? When you look in the mirror, are you happy with what you see?
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is slightly warm, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as sexual lust; style: storytelling, conversational; average recording, quiet background; mildly explicit content; genuineness 2.5/6; vocal-burst blend 3.2/10; 6.8s, EN.
941719_00006616 · in -20.6 dBFS · gain +0.6 dB · podcast-04973
(fatigue exhaustion, confusion, shame·slow, very low-energy, relaxed, casual)So I'm gonna just leave that to y'all while I actually get my thoughts together, and before I start going, because Lord knows, (ahem) yeah. So just think about that.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is slightly cool, slightly dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as fatigue exhaustion, confusion, shame; style: casual, conversational; below-average recording, quiet background; mildly explicit content; genuineness 5.5/6; vocal-burst blend 2.1/10; 16.0s, EN.
941719_00007292 · in -21.6 dBFS · gain +1.6 dB · podcast-00074
(doubt, disappointment, confusion ·normal-paced, subdued, neutral tension, casual)Okay, so solitude or standing in it, and I felt like that was a very strong word to choose, standing in solitude. Because many people think to be alone is to be weak, when actually is the complete opposite, as far as my opinion. Basically, this whole podcast is my opinion, so I'm not gonna keep saying that, but you know. So, in my opinion, I feel like a being able to be alone is a power that many people lack, if I'm just being honest,
full caption & clip details
A young adult somewhat feminine voice; delivery is subdued, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as doubt, disappointment, confusion; style: casual, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 4.3/10; 29.9s, EN.
941719_00008888 · in -22.3 dBFS · gain +2.3 dB · podcast-04962
This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Contemplation around average — 0.58, higher than 58 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.27.
At the same time Sourness goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.62 (higher than 62 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.06 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.18 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.20 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.18, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 24 s · de · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.185 before conversion and 0.839 after — it rose by 0.654. Neighbour-to-neighbour the worst pair went 0.227 → 0.839. (The earlier render, with segment 1 left raw, scores 0.754 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.268 in the original and +0.686 after conversion — 256 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Sourness, -0.382 became -0.369.
Quality. Mean predicted overall quality across the segments went 2.87 → 3.04 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.185 → 0.839+0.654identity cos neighbours 0.227 → 0.839d_b rescored +0.268 → +0.686d_a rescored -0.382 → -0.369d_a mined -0.358d_b mined 0.269min_cos_consec (site) 0.1995min_cos_anchor (site) 0.1768dataset podcastlang despeaker 564034total 23.1schain gain +1.0 dBseam step 2.2 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(sourness, amusement, jealousy and envy · fast, some disfluency, average clarity, monologue)aber so kleine Märkte, oder kleine, kleine Läden, private Schmucklinien oder was weiß ich.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as sourness, amusement, jealousy and envy; style: monologue, narration; good recording, no background noise; genuineness 2.9/6; vocal-burst blend 0.9/10; 5.4s, DE.
564034_00185888 · in -38.2 dBFS · gain +18.2 dB · podcast-01839
(confusion, intoxication altered states of consciousness, anger·normal-paced, frequent disfluency, somewhat unclear, casual)ja komplett verwüstet einfach mit (low mumble) zivilen Autos und und und das ist immer das, was mich stört. Und (ahem) das ist jetzt nicht nur mit der
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as confusion, intoxication altered states of consciousness, anger; style: casual, monologue; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 0.0/10; 8.2s, DE.
564034_00187024 · in -33.5 dBFS · gain +13.5 dB · podcast-01729
(measured, frequent disfluency, somewhat unclear, monologue)mit der Bewegung, die da gerade in Amerika stattfindet, sondern es ist auch, (low mumble) also oftmals ja auch hier in Deutschland, dass (ahem) zu diversen Themen (ahem) Demos stattfinden.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 0.3/10; 10.0s, DE.
564034_00187856 · in -34.4 dBFS · gain +14.4 dB · podcast-01725
This chain comes from the proxy rule: the same two-sided test as above, but because Hope Enthusiasm Optimism is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Hope Enthusiasm Optimism around average — 0.57, higher than 57 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.35.
At the same time Teasing goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.63. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.13, then -0.01 — not a clean run: step 3 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.37 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.37 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.37, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 49 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.251 before conversion and 0.505 after — it rose by 0.254. Neighbour-to-neighbour the worst pair went 0.285 → 0.642. (The earlier render, with segment 1 left raw, scores 0.422 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.367 in the original and +0.044 after conversion — 12 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Teasing, -0.633 became -0.175.
Quality. Mean predicted overall quality across the segments went 2.75 → 2.96 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.251 → 0.505+0.254identity cos neighbours 0.285 → 0.642d_b rescored +0.367 → +0.044d_a rescored -0.633 → -0.175d_a mined -0.633d_b mined 0.352min_cos_consec (site) 0.3665min_cos_anchor (site) 0.3665dataset podcastlang enspeaker 255928total 48.5schain gain +2.7 dBseam step 6.5 dBcrossfades 100/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, moderately variable, some disfluency, average clarity
(teasing, amusement, sexual lust · brisk, energised, neutral tension, casual)uglier? People are gonna be like take the pull who was ugly, who had the bigger who's uglier?
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as teasing, amusement, sexual lust; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 5.8/6; vocal-burst blend 5.6/10; 5.0s, EN.
255928_00323392 · in -23.5 dBFS · gain +3.5 dB · podcast-00141
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as embarrassment, amusement, shame; style: conversational, playful; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 3.7/10; 14.2s, EN.
255928_00323896 · in -26.3 dBFS · gain +6.3 dB · podcast-03607
(infatuation, contemplation, pain·brisk, normally alert, slightly relaxed, casual)You know, like where do you feel like you're following your path? Keep following it, and people are going to be attracted to you for that, right? And so, like,
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as infatuation, contemplation, pain; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 10.0/10; 22.0s, EN.
255928_00325316 · in -30.3 dBFS · gain +10.3 dB · podcast-06211
(hope enthusiasm optimism· brisk, energised, neutral tension, conversational)if that's a wound for you, and on top of it, now you are an attractive person out in this world. It's so hard out here being pretty.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism; style: conversational, casual; good recording, no background noise; genuineness 2.4/6; vocal-burst blend 3.3/10; 7.8s, EN.
255928_00327568 · in -27.5 dBFS · gain +7.5 dB · podcast-06196
This chain comes from the proxy rule: the same two-sided test as above, but because Embarrassment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Embarrassment around average — 0.55, higher than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.41.
At the same time Sexual Lust goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.16, then +0.05, then -0.01 — not a clean run: step 4 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.13 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.16 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.13, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 52 s · en · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.149 before conversion and 0.643 after — it rose by 0.493. Neighbour-to-neighbour the worst pair went 0.143 → 0.635. (The earlier render, with segment 1 left raw, scores 0.519 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.413 in the original and +0.302 after conversion — 73 % of the delta retained, which is most of it. On the other named axis, Sexual Lust, -0.273 became -0.482.
Quality. Mean predicted overall quality across the segments went 2.85 → 3.20 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.149 → 0.643+0.493identity cos neighbours 0.143 → 0.635d_b rescored +0.413 → +0.302d_a rescored -0.273 → -0.482d_a mined -0.271d_b mined 0.413min_cos_consec (site) 0.1636min_cos_anchor (site) 0.1329dataset podcastlang enspeaker 134645total 51.2schain gain +3.6 dBseam step 2.0 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, balanced body
(sexual lust, teasing, amusement · brisk, energised, slightly tense, casual)you take out TMAT and you're basically at that point, you're good to go and take on the final boss. But there's other shit here you can do too.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as sexual lust, teasing, amusement; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.8/6; vocal-burst blend 4.5/10; 6.6s, EN.
134645_00573232 · in -19.6 dBFS · gain -0.4 dB · podcast-03591
(infatuation, fear·normal-paced, normally alert, slightly relaxed, casual)(low mumble) Uh most notably there is a an organ that if you press all of the keys on the controller outside of (low mumble) uh the D pad and start and select, at the exact same time it unlocks a gate
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as infatuation, fear; style: casual, playful; good recording, quiet background; genuineness 4.1/6; vocal-burst blend 2.4/10; 14.0s, EN.
134645_00573888 · in -28.1 dBFS · gain +8.1 dB · podcast-03623
(pride, confusion, amusement· normal-paced, normally alert, slightly relaxed, casual)that you can walk through to pick up an item. But I could never pull it off 'cause I can't put I apparently don't have the the coordination and skills to press every single button on my controller at the exact same
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as pride, confusion, amusement; style: casual, authoritative; good recording, no background noise; genuineness 3.5/6; vocal-burst blend 3.8/10; 10.5s, EN.
134645_00575280 · in -32.3 dBFS · gain +12.3 dB · podcast-03615
(embarrassment, amusement, astonishment surprise· normal-paced, normally alert, neutral tension, casual)And Jake didn't even try. (contented sigh) Yeah. I
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as embarrassment, amusement, astonishment surprise; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.1/6; vocal-burst blend 3.7/10; 14.4s, EN.
134645_00576544 · in -22.9 dBFS · gain +2.9 dB · podcast-03871
(embarrassment, shame· normal-paced, normally alert, neutral tension, casual)Right. But yeah, at that end I did draw Eden from (low mumble) uh team it. I know we probably said that earlier, but
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as embarrassment, shame; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 2.3/10; 6.5s, EN.
134645_00578767 · in -21.7 dBFS · gain +1.7 dB · podcast-03882
This chain comes from the proxy rule: the same two-sided test as above, but because Pleasure Ecstasy is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Pleasure Ecstasy clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.26.
At the same time Impatience and Irritability goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.06 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.60 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.71 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.60, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 13 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.644 before conversion and 0.663 after — it rose by 0.019. Neighbour-to-neighbour the worst pair went 0.658 → 0.663. (The earlier render, with segment 1 left raw, scores 0.472 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pleasure Ecstasy moved +0.257 in the original and +0.161 after conversion — 63 % of the delta retained. On the other named axis, Impatience and Irritability, -0.268 became -0.332.
Quality. Mean predicted overall quality across the segments went 2.51 → 2.77 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.644 → 0.663+0.019identity cos neighbours 0.658 → 0.663d_b rescored +0.257 → +0.161d_a rescored -0.268 → -0.332d_a mined -0.267d_b mined 0.257min_cos_consec (site) 0.7139min_cos_anchor (site) 0.6023dataset podcastlang enspeaker 313050total 12.7schain gain +0.5 dBseam step 3.4 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, average recording, normal-paced, normally alert, some disfluency, average clarity, light breath
(impatience and irritability, intoxication altered states of consciousness, disgust · slightly relaxed, fairly steady, moderate pitch range, casual)a lot of he made a whole front. That motherfucker knows how to start a franchise, I'll tell
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as impatience and irritability, intoxication altered states of consciousness, disgust; style: casual, conversational; average recording, no background noise; genuineness 5.7/6; vocal-burst blend 3.1/10; 3.5s, EN.
313050_00290496 · in -22.1 dBFS · gain +2.1 dB · podcast-03839
(amusement, sexual lust, disgust ·neutral tension, moderately variable, wide pitch range, casual)surprised that no one's like started another one off Malignant. I loved Malignant.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as amusement, sexual lust, disgust; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.2/6; vocal-burst blend 3.1/10; 5.7s, EN.
313050_00291136 · in -19.3 dBFS · gain -0.7 dB · podcast-03843
(pleasure ecstasy, elation, affection·fully relaxed, moderately variable, wide pitch range, casual)That was so much fucking fun. I loved that movie.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, moderately variable; timbre is neutral-toned, dark, very rough, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as pleasure ecstasy, elation, affection; style: casual, conversational; average recording, quiet background; explicit content; genuineness 3.3/6; vocal-burst blend 3.4/10; 3.9s, EN.
313050_00292904 · in -22.0 dBFS · gain +2.0 dB · podcast-03862
This chain comes from the proxy rule: the same two-sided test as above, but because Embarrassment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Embarrassment clearly present — 0.73, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.26.
At the same time Relief goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.56 (higher than 56 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.15, then +0.11, then -0.01 — not a clean run: step 3 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.26 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.26 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.26, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 48 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.226 before conversion and 0.618 after — it rose by 0.392. Neighbour-to-neighbour the worst pair went 0.027 → 0.415. (The earlier render, with segment 1 left raw, scores 0.530 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.253 in the original and +0.163 after conversion — 64 % of the delta retained. On the other named axis, Relief, -0.413 became -0.415.
Quality. Mean predicted overall quality across the segments went 2.73 → 2.90 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.226 → 0.618+0.392identity cos neighbours 0.027 → 0.415d_b rescored +0.253 → +0.163d_a rescored -0.413 → -0.415d_a mined -0.414d_b mined 0.255min_cos_consec (site) 0.2562min_cos_anchor (site) 0.2634dataset podcastlang enspeaker 583280total 46.7schain gain +4.9 dBseam step 0.9 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · fairly steady
(relief, triumph, pride · measured, very low-energy, relaxed, whispered)And so I I got I got three or four years there, um, (contented sigh) which was awesome. So now I had I had distributed
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is slightly warm, slightly dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, neutral openness; reads as relief, triumph, pride; style: whispered, monologue; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 3.5/10; 6.0s, EN.
583280_00067008 · in -47.1 dBFS · gain +27.1 dB · podcast-02049
(contemplation, longing, shame·slow, very low-energy, relaxed, monologue)you know, and now I had had experience with the with the manufacturer.
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly warm, dark, rough, very full; slurred, some disfluency, fairly narrow pitch, audible breath; affect is mildly negative, submissive, neutral openness; reads as contemplation, longing, shame; style: monologue, whispered; average recording, no background noise; genuineness 2.8/6; vocal-burst blend 6.2/10; 3.1s, EN.
583280_00067759 · in -48.6 dBFS · gain +28.6 dB · podcast-02052
(embarrassment, shame, intoxication altered states of consciousness·normal-paced, very low-energy, neutral tension, casual)I mean that I think hearing you know, as I've been learning about the industry, I think like I'm I'm ticking them off like in my mind, I'm like, Man, you've you've you've gone learned the basics, come up from the beginning in three of the five really three of the five categories of the plumbing industry.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as embarrassment, shame, intoxication altered states of consciousness; style: casual, whispered; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 10.0/10; 23.7s, EN.
583280_00074380 · in -44.8 dBFS · gain +24.8 dB · podcast-05708
(embarrassment, contemplation, intoxication altered states of consciousness · normal-paced, normally alert, neutral tension, casual)I mean, that's uh that's like a three star general right (chuckle) there. You know what I mean? Like, you know, I mean, the only thing the only thing, you know, that is not in there is, you know, actually, you know, either being in the field or being in like facilities maintenance, you know,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as embarrassment, contemplation, intoxication altered states of consciousness; style: casual, conversational; good recording, quiet background; genuineness 6.0/6; vocal-burst blend 7.8/10; 14.5s, EN.
583280_00076800 · in -42.5 dBFS · gain +22.5 dB · podcast-02052
This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Interest clearly present — 0.59, higher than 59 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.39.
At the same time Longing goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.18, then +0.21 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.14 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.14 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.14, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 45 s · de · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.199 before conversion and 0.754 after — it rose by 0.555. Neighbour-to-neighbour the worst pair went 0.236 → 0.819. (The earlier render, with segment 1 left raw, scores 0.730 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.395 in the original and +0.538 after conversion — 136 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Longing, -0.298 became -0.337.
Quality. Mean predicted overall quality across the segments went 2.92 → 3.28 (+0.36) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.199 → 0.754+0.555identity cos neighbours 0.236 → 0.819d_b rescored +0.395 → +0.538d_a rescored -0.298 → -0.337d_a mined -0.303d_b mined 0.393min_cos_consec (site) 0.1434min_cos_anchor (site) 0.1405dataset podcastlang despeaker 410357total 44.4schain gain +2.0 dBseam step 0.9 dBcrossfades 150/100 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, good recording, quiet background, normally alert, slightly relaxed, fairly steady
(longing, fear · brisk, casual, conversational)du so eine Flugshow machst oder so und die sind noch irgendwie ein Kilometer weg und machen über dem See in der Nähe dann, ne, du stehst am Boden und die
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing, fear; style: casual, conversational; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 3.8/10; 6.9s, DE.
410357_00590136 · in -17.4 dBFS · gain -2.6 dB · podcast-04632
(confusion, sexual lust, pleasure ecstasy·normal-paced, casual, conversational)(ahem) machen so ihre Flugshowmanöver. Also das muss so laut sein, also deine Hosen flattern vom Bass und die sind halt noch Kilometer weg. Also es ist irgendwie.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, sexual lust, pleasure ecstasy; style: casual, conversational; good recording, quiet background; genuineness 4.1/6; vocal-burst blend 2.8/10; 8.1s, DE.
410357_00590824 · in -18.1 dBFS · gain -1.9 dB · podcast-04639
(interest, jealousy and envy, astonishment surprise· normal-paced, casual, conversational)Da (low mumble) passiert schon mal 20 Jahren.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, jealousy and envy, astonishment surprise; style: casual, conversational; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 3.3/10; 29.7s, DE.
410357_00591800 · in -19.3 dBFS · gain -0.7 dB · podcast-01660
This chain comes from the proxy rule: the same two-sided test as above, but because Doubt is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Doubt around average — 0.49, lower than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.45.
At the same time Sexual Lust goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.07, then +0.16, then +0.23 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.23 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.21 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.23, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 58 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.230 before conversion and 0.732 after — it rose by 0.502. Neighbour-to-neighbour the worst pair went 0.202 → 0.621. (The earlier render, with segment 1 left raw, scores 0.616 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.453 in the original and +0.434 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Sexual Lust, -0.267 became -0.104.
Quality. Mean predicted overall quality across the segments went 2.78 → 3.16 (+0.38) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.230 → 0.732+0.502identity cos neighbours 0.202 → 0.621d_b rescored +0.453 → +0.434d_a rescored -0.267 → -0.104d_a mined -0.282d_b mined 0.454min_cos_consec (site) 0.2094min_cos_anchor (site) 0.2330dataset podcastlang enspeaker 670826total 57.3schain gain +2.3 dBseam step 1.2 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · slightly relaxed
(sexual lust, contemplation, infatuation · slow, very low-energy, steady, whispered)really gives her full attention and really sits with and absorbs what the person is saying and doesn't rush to a response.
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is slightly warm, slightly dark, slightly rough, full; somewhat unclear, some disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, neutral openness; reads as sexual lust, contemplation, infatuation; style: whispered, monologue; average recording, no background noise; genuineness 1.8/6; vocal-burst blend 1.1/10; 7.8s, EN.
670826_00015567 · in -30.8 dBFS · gain +10.8 dB · podcast-04619
(infatuation, contentment, affection·normal-paced, subdued, fairly steady, casual)Right. You've had that experience, right? We all have, you know, of having somebody that's with you but not really with you, like, you know, engaged in a conversation with you but not really listening. So, you know, especially when we get into matters of the heart, you know, things that are really painful or whatever, we really want to be able to be fully present to them. So to make make the space for that to happen, to
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as infatuation, contentment, affection; style: casual, whispered; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 9.1/10; 22.4s, EN.
670826_00016383 · in -29.6 dBFS · gain +9.6 dB · podcast-02977
(contemplation·measured, very low-energy, fairly steady, whispered)hold that space, right? Right. Cause when you actually slow yourself down as you're listening, and they begin to share, well, then your your friend is sharing, and then if you don't leap in with something too quick, they may actually come up with the next sentence that helps them resolve what they were already saying.
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, slightly guarded; reads as contemplation; style: whispered, didactic; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 2.6/10; 22.5s, EN.
670826_00018623 · in -31.3 dBFS · gain +11.3 dB · podcast-06246
(doubt, teasing, sourness·normal-paced, normally alert, fairly steady, casual)will totally own that and they they will they you might actually find yourself saying very little.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as doubt, teasing, sourness; style: casual, conversational; good recording, quiet background; genuineness 3.7/6; vocal-burst blend 3.8/10; 5.2s, EN.
670826_00021248 · in -28.7 dBFS · gain +8.7 dB · podcast-04620
This chain comes from the proxy rule: the same two-sided test as above, but because Embarrassment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Embarrassment clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.30.
At the same time Doubt goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.09, then -0.12, then +0.13 — not a clean run: step 3 moves back the other way by 0.12 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.17 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.24 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.17, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 96 s · en · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.212 before conversion and 0.623 after — it rose by 0.411. Neighbour-to-neighbour the worst pair went 0.275 → 0.830. (The earlier render, with segment 1 left raw, scores 0.561 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.308 in the original and +0.294 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Doubt, -0.301 became -0.238.
Quality. Mean predicted overall quality across the segments went 3.17 → 3.28 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.212 → 0.623+0.411identity cos neighbours 0.275 → 0.830d_b rescored +0.308 → +0.294d_a rescored -0.301 → -0.238d_a mined -0.301d_b mined 0.303min_cos_consec (site) 0.2365min_cos_anchor (site) 0.1673dataset podcastlang enspeaker 904575total 94.4schain gain +3.6 dBseam step 0.5 dBcrossfades 100/150/150/100 ms
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, quiet background, normal-paced, normally alert, fairly steady, moderate pitch range
(doubt, confusion, contemplation · slightly relaxed, some disfluency, average clarity, conversational)So, you know, a lot of it is kind of (low mumble) a hustle and fake it till you make it. (low mumble) Um is was it the same for you with real estate? Like how do you how do you so how do you jump into doing real estate with no background in it?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as doubt, confusion, contemplation; style: conversational, casual; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.7/10; 10.8s, EN.
904575_00045272 · in -21.0 dBFS · gain +1.0 dB · podcast-03856
(fatigue exhaustion, intoxication altered states of consciousness, contentment·relaxed, frequent disfluency, average clarity, casual)(low mumble) Um probably like a lot of people, you know, thinking it's time to buy your first home. You've rented for so long and and (low mumble) uh and and it's kind of time to pull the trigger and you have that thought in your mind, oh, I'm just throwing this rent money away, you know, I should put it into something I'm gonna own.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as fatigue exhaustion, intoxication altered states of consciousness, contentment; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 9.1/10; 15.2s, EN.
904575_00046472 · in -20.4 dBFS · gain +0.4 dB · podcast-06326
(shame, embarrassment, relief·slightly relaxed, some disfluency, average clarity, casual)(low mumble) Um it and for me, it started a little before that. (low mumble) Uh when I was in Maine, I had a a friend of mine who uh (low mumble) who I I mentioned I ran track, and so there was a runner a few years older than I was, one of the best runners in the state. And and years later, (low mumble) uh at really out of college, we reconnected and and and formed a friendship.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as shame, embarrassment, relief; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 8.4/10; 18.8s, EN.
904575_00048040 · in -20.4 dBFS · gain +0.4 dB · podcast-02487
(awe, infatuation, astonishment surprise·neutral tension, some disfluency, average clarity, casual)(low mumble) Uh and at that point, he was pretty heavily into real estate and then had to build up a pretty decent portfolio. And I I I don't even think I realized to what extent at that time. (low mumble) Um, but had a really cool story. He, you know, had a college loan and used that money to get his first duplex and (low mumble) uh and then kind of that got him started, right? And so that that spiked my initial interest in in real estate and
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as awe, infatuation, astonishment surprise; style: casual, conversational; good recording, quiet background; genuineness 4.4/6; vocal-burst blend 9.5/10; 22.3s, EN.
904575_00049920 · in -20.6 dBFS · gain +0.6 dB · podcast-06328
(embarrassment, shame, relief·slightly relaxed, some disfluency, somewhat unclear, casual)(ahem) that (low mumble) was pretty that (low mumble) was (low mumble) really early on (low mumble) then.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as embarrassment, shame, relief; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 9.6/10; 27.9s, EN.
904575_00052200 · in -20.6 dBFS · gain +0.6 dB · podcast-06326
This chain comes from the proxy rule: the same two-sided test as above, but because Embarrassment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Embarrassment clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.33.
At the same time Concentration goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.20 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 32 s · nl · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.873 before conversion and 0.844 after — it fell by 0.030. Neighbour-to-neighbour the worst pair went 0.798 → 0.783. (The earlier render, with segment 1 left raw, scores 0.786 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.334 in the original and +0.287 after conversion — 86 % of the delta retained, which is most of it. On the other named axis, Concentration, -0.262 became -0.472.
Quality. Mean predicted overall quality across the segments went 3.04 → 3.17 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.873 → 0.844-0.030identity cos neighbours 0.798 → 0.783d_b rescored +0.334 → +0.287d_a rescored -0.262 → -0.472d_a mined -0.265d_b mined 0.334min_cos_consec (site) 0.8585min_cos_anchor (site) 0.8729dataset podcastlang nlspeaker 675790total 31.8schain gain +2.9 dBseam step 2.5 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, good recording, quiet background, normally alert, some disfluency, light breath
(concentration, astonishment surprise, interest · brisk, slightly relaxed, fairly steady, dramatic)En ik ben er ook achter gekomen dat de manier waarop ik liefde ontvang, of wil ontvangen, en de manier waarop ik liefde geef, ook heel anders is. Maar ik zal ze misschien eerst een keer alle vijf gaan overlopen. Je hebt words of affirmation.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as concentration, astonishment surprise, interest; style: dramatic, authoritative; good recording, quiet background; genuineness 1.4/6; vocal-burst blend 1.7/10; 13.3s, NL.
675790_00163300 · in -28.4 dBFS · gain +8.4 dB · podcast-04214
(shame, doubt·normal-paced, slightly relaxed, fairly steady, monologue)Words of affirmation, dat wil zeggen dat je heel graag zo bevestiging krijgt. Dat je dat echt verbaal ook moet gaan horen van, schat, ik ben vier op je of, schat, bedankt dat jij vandaag hebt gekookt.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as shame, doubt; style: monologue, authoritative; good recording, quiet background; genuineness 1.0/6; vocal-burst blend 0.0/10; 12.4s, NL.
675790_00164632 · in -28.9 dBFS · gain +8.9 dB · podcast-04214
(embarrassment·brisk, neutral tension, moderately variable, casual)In deze context mogen we wel koosnaapjes gebruiken, want het gaat over love. We mogen wel Popemia schat zeggen.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as embarrassment; style: casual, conversational; good recording, quiet background; genuineness 4.7/6; vocal-burst blend 5.3/10; 6.4s, NL.
675790_00165868 · in -23.9 dBFS · gain +3.9 dB · podcast-04214