This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_c-podcast-B1.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words
are a real non-speech sound, printed where it happens.
The full generated caption for any clip is under “full caption & clip
details”. Its perceived-gender and background-noise clauses were re-rendered from
the numeric buckets, because the versions stored in the corpus index had those two
ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
The tags and captions describe the ORIGINAL clips — they are the
corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the
converted audio are the ones printed in each card's own paragraph and chips.
This chain comes from the one-sided rule: only Sourness had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Sourness clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.30.
Nothing was asked of the other axis, and in fact Affection drifts down from 0.96 to 0.60 (-0.36), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.14 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.53 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.62 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.53, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 54 s · es · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.574 before conversion and 0.716 after — it rose by 0.143. Neighbour-to-neighbour the worst pair went 0.652 → 0.714. (The earlier render, with segment 1 left raw, scores 0.547 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sourness moved +0.298 in the original and +0.852 after conversion — 286 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Affection, -0.359 became -0.196.
Quality. Mean predicted overall quality across the segments went 2.73 → 3.02 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.574 → 0.716+0.143identity cos neighbours 0.652 → 0.714d_b rescored +0.298 → +0.852d_a rescored -0.359 → -0.196d_a mined -0.360d_b mined 0.298min_cos_consec (site) 0.6211min_cos_anchor (site) 0.5258dataset podcastlang esspeaker 907370total 53.4schain gain +1.3 dBseam step 2.4 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, quiet background, normally alert, neutral tension, some disfluency, somewhat unclear
(affection, triumph, embarrassment · normal-paced, fairly steady, moderate pitch range, casual)Y en el filial rojiblanco allí en mareo con Abelardo, vaya dos mitos que te dirigieron allí en el (low mumble) Sporting.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as affection, triumph, embarrassment; style: casual, conversational; below-average recording, quiet background; genuineness 4.7/6; vocal-burst blend 10.0/10; 20.8s, ES.
907370_00008616 · in -18.8 dBFS · gain -1.2 dB · podcast-00488
(affection, relief, jealousy and envy· normal-paced, moderately variable, moderate pitch range, casual)(low mumble) El Racing parece que ha despejado dudas con el tres cero al Málaga y el Sporting que empezó muy bien, cinco derrotas consecutivas, en una segunda división que es de locos,
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as affection, relief, jealousy and envy; style: casual, conversational; below-average recording, quiet background; genuineness 4.7/6; vocal-burst blend 9.3/10; 23.6s, ES.
907370_00010692 · in -18.7 dBFS · gain -1.3 dB · podcast-01169
(sourness, confusion, embarrassment·fast, moderately variable, wide pitch range, casual)pero evidentemente, pues nuevo entrenador, a ver qué pasa, pero la gente como que se ha quedado un poco fría, ¿no? por esas cinco derrotas consecutivas. Si
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as sourness, confusion, embarrassment; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.4/6; vocal-burst blend 10.0/10; 9.3s, ES.
907370_00013054 · in -19.1 dBFS · gain -0.9 dB · podcast-01171
This chain comes from the one-sided rule: only Thankfulness Gratitude had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Thankfulness Gratitude clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.29.
Nothing was asked of the other axis, and in fact Infatuation barely moves at all, sitting near 0.98 throughout.
It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.05 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 40 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.842 before conversion and 0.760 after — it fell by 0.082. Neighbour-to-neighbour the worst pair went 0.848 → 0.760. (The earlier render, with segment 1 left raw, scores 0.565 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.290 in the original and +0.184 after conversion — 63 % of the delta retained. On the other named axis, Infatuation, -0.028 became +0.054.
Quality. Mean predicted overall quality across the segments went 2.93 → 3.07 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.842 → 0.760-0.082identity cos neighbours 0.848 → 0.760d_b rescored +0.290 → +0.184d_a rescored -0.028 → +0.054d_a mined -0.021d_b mined 0.291min_cos_consec (site) 0.8555min_cos_anchor (site) 0.8391dataset podcastlang enspeaker 882347total 39.3schain gain +2.6 dBseam step 0.7 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, moderately variable
(infatuation, embarrassment, shame · measured, normally alert, slightly relaxed, casual)around. He taught uh (wistful sigh) I think at like a maybe the middle school or high school. (ahem) Um and then he moved to Title I reading when I was young.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as infatuation, embarrassment, shame; style: casual, monologue; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 0.0/10; 7.7s, EN.
882347_00126164 · in -37.9 dBFS · gain +17.9 dB · podcast-02088
(disappointment, amusement, shame ·normal-paced, very low-energy, neutral tension, casual)Did that move to like fifth grade, fourth grade reading, I think it was. Still title one there, and then he ended his career by teaching eighth grade social studies. I mean, that's what made me a reader, that's what made me want to be a teacher. I spent my summers with him. I uh begged him to go to teacher, like go to work with your parent day. You know what (breathy giggle) I mean? Begged him, begged him, begged him.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, minimal breath; affect is mildly positive, slightly submissive, neutral openness; reads as disappointment, amusement, shame; style: casual, ASMR; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 7.1/10; 22.8s, EN.
882347_00126935 · in -37.4 dBFS · gain +17.4 dB · podcast-01349
(thankfulness gratitude, longing, relief· normal-paced, normally alert, slightly relaxed, casual)He would not take me. And then when I was in like fifth grade, finally, he told me when I was older than the kids he had in class, he would let me go with him.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as thankfulness gratitude, longing, relief; style: casual, storytelling; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 0.0/10; 9.1s, EN.
882347_00129216 · in -36.3 dBFS · gain +16.3 dB · podcast-02112
This chain comes from the one-sided rule: only Contentment had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Contentment clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.29.
Nothing was asked of the other axis, and in fact Pride barely moves at all, sitting near 0.93 throughout.
It takes 5 clips to get there. Clip to clip the moves are -0.07, then +0.17, then +0.05, then +0.13 — not a clean run: step 1 moves back the other way by 0.07 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.26 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.37 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.26, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 47 s · en · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.252 before conversion and 0.785 after — it rose by 0.533. Neighbour-to-neighbour the worst pair went 0.312 → 0.781. (The earlier render, with segment 1 left raw, scores 0.672 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.289 in the original and +0.221 after conversion — 76 % of the delta retained, which is most of it. On the other named axis, Pride, -0.012 became -0.100.
Quality. Mean predicted overall quality across the segments went 2.81 → 2.98 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.252 → 0.785+0.533identity cos neighbours 0.312 → 0.781d_b rescored +0.289 → +0.221d_a rescored -0.012 → -0.100d_a mined -0.012d_b mined 0.289min_cos_consec (site) 0.3668min_cos_anchor (site) 0.2609dataset podcastlang enspeaker 654418total 45.3schain gain +0.2 dBseam step 1.6 dBcrossfades 150/150/100/100 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, normally alert, some disfluency, average clarity
(pride · normal-paced, neutral tension, moderately variable, casual)sit out to eye, right? Like we live in one of the most segregated cities in the country. Our literally our public areas are set up in certain ways so that homeless folks can't
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as pride; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.8/6; vocal-burst blend 8.8/10; 8.2s, EN.
654418_00193280 · in -27.6 dBFS · gain +7.6 dB · podcast-02018
(amusement, teasing· normal-paced, slightly relaxed, fairly steady, casual)property. And so this is a place where we say, like, hey, all those walls are like bullshit. Yeah.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as amusement, teasing; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.2/6; vocal-burst blend 5.5/10; 4.3s, EN.
654418_00194328 · in -29.4 dBFS · gain +9.4 dB · podcast-02001
(normal-paced, neutral tension, moderately variable, casual)Like here is where everybody comes to the table. We all sit down, we all eat. Sure, there's power differentials. There always is, right?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 6.8/10; 6.6s, EN.
654418_00194768 · in -25.2 dBFS · gain +5.2 dB · podcast-02006
(infatuation, sexual lust· normal-paced, neutral tension, moderately variable, casual)Because of the socioeconomic status, et cetera. But we try our hardest to say, like, those aren't gonna matter here in the same way they matter in other places. (ahem) Um, because they really don't matter. Rich, I love this so much. This is such a great
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as infatuation, sexual lust; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.7/6; vocal-burst blend 6.4/10; 10.8s, EN.
654418_00195440 · in -29.5 dBFS · gain +9.5 dB · podcast-02002
(contentment, elation, pleasure ecstasy·brisk, neutral tension, moderately variable, casual)conversation. I love what you said. Those walls are bullshit. Like at the table, the best we can. We know we can't eliminate all the barriers, right? And like it's not everyone's perfectly equal, but the best we can, everyone is equal and valued and loved and seen at this table, right? That's right.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as contentment, elation, pleasure ecstasy; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 6.5/10; 16.0s, EN.
654418_00196680 · in -28.0 dBFS · gain +8.0 dB · podcast-02767
This chain comes from the one-sided rule: only Astonishment Surprise had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Astonishment Surprise strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.20.
Nothing was asked of the other axis, and in fact Doubt barely moves at all, sitting near 1.00 throughout.
It takes 2 clips to get there. Clip to clip the moves are +0.20 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.51 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.51 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.51, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 15 s · en · podcast
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.153 before conversion and 0.383 after — it rose by 0.230. Neighbour-to-neighbour the worst pair went 0.153 → 0.383. (The earlier render, with segment 1 left raw, scores 0.246 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.202 in the original and +0.616 after conversion — 304 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Doubt, -0.033 became -0.028.
Quality. Mean predicted overall quality across the segments went 2.25 → 2.92 (+0.67) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.153 → 0.383+0.230identity cos neighbours 0.153 → 0.383d_b rescored +0.202 → +0.616d_a rescored -0.033 → -0.028d_a mined -0.033d_b mined 0.201min_cos_consec (site) 0.5099min_cos_anchor (site) 0.5099dataset podcastlang enspeaker 662687total 14.8schain gain +2.6 dBseam step 1.4 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · fairly smooth, normal-paced, normally alert, moderately variable, some disfluency, wide pitch range
(doubt, contemplation, confusion · neutral tension, average clarity, light breath, casual)I don't know. I just am like do the well, maybe I don't know enough about the situation, but my thought is like if people are in remote areas and they're not getting like are they getting education? Like
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as doubt, contemplation, confusion; style: casual, conversational; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 8.0/10; 11.9s, EN.
662687_00164928 · in -32.4 dBFS · gain +12.4 dB · podcast-02567
(astonishment surprise, doubt, impatience and irritability·fully relaxed, slurred, heavy breath, casual)can they even take it? Oh, that's a fair question.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, fully relaxed, moderately variable; timbre is slightly cool, dark, fairly smooth, thin; slurred, some disfluency, wide pitch range, heavy breath; affect is positive, slightly submissive, neutral openness; reads as astonishment surprise, doubt, impatience and irritability; style: casual, playful; below-average recording, some background noise; genuineness 4.6/6; vocal-burst blend 4.0/10; 3.0s, EN.
662687_00166128 · in -34.5 dBFS · gain +14.5 dB · podcast-00664
This chain comes from the one-sided rule: only Contentment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Contentment strongly present — 0.76, higher than 76 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.21.
Nothing was asked of the other axis, and in fact Hope Enthusiasm Optimism drifts down from 0.97 to 0.82 (-0.15), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are -0.16, then +0.25, then +0.13 — not a clean run: step 1 moves back the other way by 0.16 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.66 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.66 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.66, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 34 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.697 before conversion and 0.792 after — it rose by 0.095. Neighbour-to-neighbour the worst pair went 0.755 → 0.614. (The earlier render, with segment 1 left raw, scores 0.684 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.212 in the original and +0.266 after conversion — 125 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Hope Enthusiasm Optimism, -0.152 became -0.116.
Quality. Mean predicted overall quality across the segments went 2.74 → 3.05 (+0.31) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.697 → 0.792+0.095identity cos neighbours 0.755 → 0.614d_b rescored +0.212 → +0.266d_a rescored -0.152 → -0.116d_a mined -0.149d_b mined 0.216min_cos_consec (site) 0.6559min_cos_anchor (site) 0.6559dataset podcastlang enspeaker 34545total 33.0schain gain +2.1 dBseam step 0.9 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · balanced body, fairly steady
(hope enthusiasm optimism, contemplation, elation · measured, normally alert, slightly relaxed, formal)Giving up on yourself. And it helps you to transform your life to actually live that life.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as hope enthusiasm optimism, contemplation, elation; style: formal, monologue; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.7/10; 5.1s, EN.
34545_00142224 · in -16.9 dBFS · gain -3.1 dB · podcast-01425
(normal-paced, normally alert, slightly relaxed, casual)at levels that that is just, you know, phenomenal.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, no background noise; genuineness 3.9/6; vocal-burst blend 2.5/10; 3.3s, EN.
34545_00143288 · in -18.7 dBFS · gain -1.3 dB · podcast-03575
(affection, hope enthusiasm optimism, elation·measured, normally alert, slightly relaxed, monologue)Beyond that, (low mumble) um, if if you finish your Dream Builder, then I take you into the life mastery system where we go for six months, um (low mumble) uh one
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, neutral openness; reads as affection, hope enthusiasm optimism, elation; style: monologue; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.0/10; 9.5s, EN.
34545_00143720 · in -17.1 dBFS · gain -2.9 dB · podcast-01432
(contentment, awe, contemplation· measured, very low-energy, relaxed, whispered)area of mastery in life per month. So these are things like your your health and well being, and we go at a much deeper level than than we could do in the 12th weeks with Dream Builder. (low mumble) Um, you know, what is abundance really
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is warm, dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as contentment, awe, contemplation; style: whispered, monologue; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 4.8/10; 15.6s, EN.
34545_00144672 · in -17.6 dBFS · gain -2.4 dB · podcast-01440
This chain comes from the one-sided rule: only Contempt had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Contempt clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.36.
Nothing was asked of the other axis, and in fact Pain drifts down from 0.99 to 0.54 (-0.45), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.16, then +0.21 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.04 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.24 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.04, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 37 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.183 before conversion and 0.463 after — it rose by 0.280. Neighbour-to-neighbour the worst pair went 0.183 → 0.463. (The earlier render, with segment 1 left raw, scores 0.412 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contempt moved +0.362 in the original and +0.334 after conversion — 92 % of the delta retained, which is essentially all of it. On the other named axis, Pain, -0.450 became -0.276.
Quality. Mean predicted overall quality across the segments went 2.62 → 2.82 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.183 → 0.463+0.280identity cos neighbours 0.183 → 0.463d_b rescored +0.362 → +0.334d_a rescored -0.450 → -0.276d_a mined -0.450d_b mined 0.362min_cos_consec (site) 0.2381min_cos_anchor (site) 0.0372dataset podcastlang enspeaker 435374total 35.9schain gain +3.8 dBseam step 0.7 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a child feminine voice · neutral-toned, fairly smooth, balanced body, average recording, normal-paced, light breath
(pain, embarrassment, fatigue exhaustion · normally alert, slightly relaxed, moderately variable, casual)I had I had to get magic to keep my eyes open. (breathy giggle)
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as pain, embarrassment, fatigue exhaustion; style: casual, playful; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 1.7/10; 3.1s, EN.
435374_00204976 · in -26.4 dBFS · gain +6.4 dB · podcast-03379
(intoxication altered states of consciousness, contemplation, interest·subdued, neutral tension, fairly steady, casual)It it it was a slow movie, yeah. There's no doubt about that. But it it it was based in reality. It's a realistic type of movie. It what didn't have all the kind of backstories and such on redemption handling stuff, it was just the man and how he was going to escape from the this prison. (low mumble) Um so it was slow, but even like in
full caption & clip details
A young adult masculine voice; delivery is subdued, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as intoxication altered states of consciousness, contemplation, interest; style: casual, conversational; average recording, quiet background; genuineness 6.0/6; vocal-burst blend 9.0/10; 20.8s, EN.
435374_00205328 · in -24.0 dBFS · gain +4.0 dB · podcast-03359
(distress, contemplation, fear·normally alert, neutral tension, fairly steady, casual)it's that that that's the kind of reality he's set up. And so this was a realistic story, and it wasn't dramatic, it was just how a person would escape that.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as distress, contemplation, fear; style: casual, conversational; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 5.1/10; 8.8s, EN.
435374_00208128 · in -25.9 dBFS · gain +5.8 dB · podcast-03377
(contempt· normally alert, slightly relaxed, fairly steady, casual)But yeah, but it is nothing nothing special about us both.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as contempt; style: casual, conversational; average recording, no background noise; genuineness 5.1/6; vocal-burst blend 0.0/10; 3.7s, EN.
435374_00209592 · in -23.8 dBFS · gain +3.8 dB · podcast-03365
This chain comes from the one-sided rule: only Relief had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Relief clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.35.
Nothing was asked of the other axis, and in fact Pain drifts down from 0.97 to 0.01 (-0.96), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.14 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.73 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.73 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.73, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 64 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.621 before conversion and 0.569 after — it fell by 0.052. Neighbour-to-neighbour the worst pair went 0.621 → 0.569. (The earlier render, with segment 1 left raw, scores 0.453 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.347 in the original and +0.347 after conversion — 100 % of the delta retained, which is essentially all of it. On the other named axis, Pain, -0.958 became -0.132.
Quality. Mean predicted overall quality across the segments went 2.79 → 3.03 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.621 → 0.569-0.052identity cos neighbours 0.621 → 0.569d_b rescored +0.347 → +0.347d_a rescored -0.958 → -0.132d_a mined -0.957d_b mined 0.347min_cos_consec (site) 0.7273min_cos_anchor (site) 0.7273dataset podcastlang enspeaker 297704total 63.2schain gain +7.0 dBseam step 2.1 dBcrossfades 100/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, average recording, quiet background, neutral tension, moderately variable
(pain, emotional numbness, embarrassment · normal-paced, normally alert, some disfluency, casual)I don't completely agree. I really like Halo Cinematics. Um (ahem) that's
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as pain, emotional numbness, embarrassment; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.7/6; vocal-burst blend 0.0/10; 4.9s, EN.
297704_00145640 · in -26.3 dBFS · gain +6.3 dB · podcast-04952
(interest, hope enthusiasm optimism, awe· normal-paced, very low-energy, frequent disfluency, casual)They don't have them out of game, but they do have them in game, and they really build on the story really well, and they look amazing. Like when the Sentinels cut through (ahem) unyielding whatever it is, (ahem) uh Atriox's ship in Halo Wars 2. Fantastic. (low mumble) Um Jerome's fight with Brutz, also amazing. Even the grave mind in Halo 2 looks great. (ahem) Especially since it was made in like 2004. And
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, hope enthusiasm optimism, awe; style: casual, monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 8.8/10; 29.8s, EN.
297704_00146176 · in -27.3 dBFS · gain +7.3 dB · podcast-04971
(relief, infatuation, jealousy and envy·measured, very low-energy, frequent disfluency, casual)but (ahem) uh yeah, like those those are really good cinematics. So I do like Riot Cinematics for League, how they're like building on the story from outside the game. That's fair. I prefer it to Overwatches. (exhausted groan) Um partly because they just seem to build a lot
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, slightly guarded; reads as relief, infatuation, jealousy and envy; style: casual, monologue; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 7.6/10; 28.7s, EN.
297704_00150495 · in -27.0 dBFS · gain +7.0 dB · podcast-04945
This chain comes from the one-sided rule: only Pride had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Pride clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.27.
Nothing was asked of the other axis, and in fact Contentment climbs from 0.84 to 0.89 (+0.05), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.09, then -0.11, then +0.24, then +0.05 — not a clean run: step 2 moves back the other way by 0.11 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.54 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.54 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.54, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 99 s · es · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.594 before conversion and 0.802 after — it rose by 0.208. Neighbour-to-neighbour the worst pair went 0.594 → 0.802. (The earlier render, with segment 1 left raw, scores 0.757 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.271 in the original and +0.420 after conversion — 155 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contentment, +0.053 became +0.052.
Quality. Mean predicted overall quality across the segments went 2.89 → 3.09 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.594 → 0.802+0.208identity cos neighbours 0.594 → 0.802d_b rescored +0.271 → +0.420d_a rescored +0.053 → +0.052d_a mined 0.054d_b mined 0.270min_cos_consec (site) 0.5409min_cos_anchor (site) 0.5409dataset podcastlang esspeaker 470731total 97.9schain gain +3.9 dBseam step 2.5 dBcrossfades 100/150/100/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-bright, fairly smooth, average clarity
(normal-paced, normally alert, neutral tension, dramatic)(low mumble) más entrada de aire, le pueden realizar más agujeros. Pero esta es la recomendación básica que hacemos para realizar la compostera. Listo, el compost ya está (ahem)
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: dramatic, monologue; average recording, no background noise; genuineness 3.8/6; vocal-burst blend 3.7/10; 11.5s, ES.
470731_00270012 · in -23.2 dBFS · gain +3.2 dB · podcast-01989
(affection, contentment, relief· normal-paced, normally alert, slightly relaxed, monologue)(ahem) para usarse para inaugurarlo luego (ahem) de dos (low mumble) días. (low mumble)
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as affection, contentment, relief; style: monologue, casual; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 8.2/10; 23.3s, ES.
470731_00271160 · in -24.5 dBFS · gain +4.5 dB · podcast-01966
(relief, disappointment·brisk, normally alert, neutral tension, casual)Y de hecho con esta basura limpia, eh, nos es más (ahem) fácil (ahem) poder separar nuestros residuos, y en el caso
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as relief, disappointment; style: casual, storytelling; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 9.1/10; 17.9s, ES.
470731_00278362 · in -25.0 dBFS · gain +5.0 dB · podcast-01975
(hope enthusiasm optimism, relief, pride· brisk, normally alert, slightly relaxed, monologue)de los plásticos, podemos hacer un ecobrique tranquilamente, (low mumble) para poder (ahem) utilizarlo en la vía construcción o donarlo a alguna organización que se encargue del tratamiento de plásticos para poder donarlo a instituciones o realizar proyectos sociales. Así que de esa manera se te va a hacer más fácil poder hacerte más consciente de tu de la generación de basura.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, relief, pride; style: monologue, dramatic; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 4.5/10; 24.1s, ES.
470731_00280152 · in -23.4 dBFS · gain +3.4 dB · podcast-01970
(pride, hope enthusiasm optimism, elation· brisk, energised, neutral tension, casual)Esperemos que les haya gustado este nuevo podcast de Despega Tu Mente. Y los esperamos en el próximo. Estén pendientes el miércoles que viene al nuevo podcast de Despega tu Mente. Entre en nuestra fanpage, acolítenos con un me gusta y podrán ver también que subimos cositas sobre la temática de los podcasts. Estén pendientes.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as pride, hope enthusiasm optimism, elation; style: casual, dramatic; below-average recording, quiet background; genuineness 4.8/6; vocal-burst blend 8.2/10; 21.7s, ES.
470731_00308016 · in -20.7 dBFS · gain +0.7 dB · podcast-01966
Emotional Numbness ↑ (unconstrained axis: Intoxication Altered States of Consciousness)identity +0.26emotion 125 % c-podcast-B1 · #9
This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Emotional Numbness around average — 0.54, higher than 54 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.36.
Nothing was asked of the other axis, and in fact Intoxication Altered States of Consciousness drifts down from 0.99 to 0.55 (-0.45), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.02, then +0.02, then +0.17, then +0.14 — a plateau around step 2, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.63 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.64 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.63, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 91 s · en · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.317 before conversion and 0.574 after — it rose by 0.257. Neighbour-to-neighbour the worst pair went 0.259 → 0.482. (The earlier render, with segment 1 left raw, scores 0.493 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.360 in the original and +0.451 after conversion — 125 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Intoxication Altered States of Consciousness, -0.429 became -0.628.
Quality. Mean predicted overall quality across the segments went 2.63 → 3.16 (+0.52) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.317 → 0.574+0.257identity cos neighbours 0.259 → 0.482d_b rescored +0.360 → +0.451d_a rescored -0.429 → -0.628d_a mined -0.447d_b mined 0.360min_cos_consec (site) 0.6353min_cos_anchor (site) 0.6301dataset podcastlang enspeaker 315471total 89.2schain gain +3.1 dBseam step 0.6 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: an elderly masculine voice · below-average recording, quiet background, relaxed, frequent disfluency
(intoxication altered states of consciousness, infatuation, pleasure ecstasy · measured, very low-energy, fairly steady, monologue)You know, I was 15 or 16 when it came out, you know, in film, and I'm like, I'm old enough to experience this. I really am, folks. I'm okay. You know, I was the kid that was carting around, you know, Stephen King, Clive Barker, (low mumble) uh David Morel, you know, all these like novels by these people that, you know,
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, dark, rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, submissive, neutral openness; reads as intoxication altered states of consciousness, infatuation, pleasure ecstasy; style: monologue, whispered; below-average recording, quiet background; genuineness 4.3/6; vocal-burst blend 7.5/10; 22.0s, EN.
315471_00161664 · in -40.8 dBFS · gain +20.8 dB · podcast-03060
(confusion, astonishment surprise, embarrassment·slow, very low-energy, steady, whispered)to me it just made sense. It fell in line with what I was into. But the community that I was in, oh no. No, no, no, no. I remember (low mumble) um first day of actual high school, not middle school, because the middle school and high school were together at a highland. (low mumble) Um
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly cool, dark, rough, slightly thin; slurred, frequent disfluency, narrow pitch range, audible breath; affect is mildly negative, submissive, neutral openness; reads as confusion, astonishment surprise, embarrassment; style: whispered, monologue; below-average recording, quiet background; genuineness 3.8/6; vocal-burst blend 3.3/10; 16.4s, EN.
315471_00163858 · in -37.8 dBFS · gain +17.8 dB · podcast-03036
(teasing, astonishment surprise, intoxication altered states of consciousness·measured, very low-energy, fairly steady, casual)I was like, Do you have any, you know, do you have any books from this author? No, can't carry it. I would love to. And I went through this list in the librarian, she's like, You're better off just going to the Holmes County Library. I was like, Yeah, that's where I get most of my books. But I did discover (ahem) um
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, submissive, neutral openness; reads as teasing, astonishment surprise, intoxication altered states of consciousness; style: casual, whispered; below-average recording, quiet background; genuineness 5.1/6; vocal-burst blend 6.7/10; 19.2s, EN.
315471_00165496 · in -41.5 dBFS · gain +21.6 dB · podcast-03038
(pride, intoxication altered states of consciousness, contentment·slow, very low-energy, steady, whispered)two authors, no, three authors in my own high school library that I still love to this day. Robert Heinlein, Eric S. Nyland, and David Morell. You know, three authors that I have honestly influenced my own writing, you know. And then it wasn't until I was out of the house and I randomly discovered a paperback novel by a damn uh
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly cool, dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, neutral openness; reads as pride, intoxication altered states of consciousness, contentment; style: whispered, monologue; below-average recording, quiet background; genuineness 3.5/6; vocal-burst blend 4.3/10; 27.6s, EN.
315471_00167416 · in -43.9 dBFS · gain +23.9 dB · podcast-03033
(slow, lethargic, steady, conversational)by a nam by a man named (ahem) uh Andrew Fox.
full caption & clip details
An elderly masculine voice; delivery is lethargic, slow, relaxed, steady; timbre is slightly warm, very dark, rough, very thin; slurred, frequent disfluency, narrow pitch range, audible breath; affect is mildly negative, submissive, neutral openness; no dominant emotion; style: conversational, casual; below-average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.4/10; 4.8s, EN.
315471_00170176 · in -40.2 dBFS · gain +20.2 dB · podcast-03036
This chain comes from the one-sided rule: only Affection had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Affection clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.34.
Nothing was asked of the other axis, and in fact Jealousy and Envy drifts down from 0.98 to 0.60 (-0.38), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.18, then +0.16 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.76 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.69 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.76, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 21 s · es · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.779 before conversion and 0.588 after — it fell by 0.191. Neighbour-to-neighbour the worst pair went 0.731 → 0.588. (The earlier render, with segment 1 left raw, scores 0.554 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Affection moved +0.346 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Jealousy and Envy, -0.379 became -0.615.
Quality. Mean predicted overall quality across the segments went 2.60 → 2.99 (+0.39) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.779 → 0.588-0.191identity cos neighbours 0.731 → 0.588d_b rescored +0.346 → +0.000d_a rescored -0.379 → -0.615d_a mined -0.379d_b mined 0.340min_cos_consec (site) 0.6931min_cos_anchor (site) 0.7573dataset podcastlang esspeaker 418145total 20.0schain gain +2.5 dBseam step 0.6 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, balanced body, normally alert
(jealousy and envy, pride, elation · normal-paced, neutral tension, moderately variable, casual)Mira, David Esteban Vargas dice, y volado lo dejó fuera, es el mejor extremo que hay. Para mí tiene que jugar volado.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as jealousy and envy, pride, elation; style: casual, conversational; average recording, some background noise; genuineness 6.0/6; vocal-burst blend 5.4/10; 7.1s, ES.
418145_00309720 · in -25.4 dBFS · gain +5.4 dB · podcast-01310
(normal-paced, slightly relaxed, moderately variable, casual)Pero la mejor jugada la hizo como extremo.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, audible breath; affect is mildly negative, slightly dominant, neutral openness; no dominant emotion; style: casual, conversational; below-average recording, quiet background; genuineness 4.5/6; vocal-burst blend 2.6/10; 3.7s, ES.
418145_00311336 · in -24.1 dBFS · gain +4.2 dB · podcast-01294
(affection·measured, relaxed, fairly steady, casual)Con Venegas, cuando encaro. ¿Sabes qué? Yo creo que Venegas le puede dar un plus (low mumble) importante a
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, very dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as affection; style: casual, conversational; below-average recording, quiet background; genuineness 5.3/6; vocal-burst blend 3.4/10; 9.6s, ES.
418145_00311700 · in -25.3 dBFS · gain +5.3 dB · podcast-01271
This chain comes from the one-sided rule: only Impatience and Irritability had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Impatience and Irritability clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.35.
Nothing was asked of the other axis, and in fact Doubt barely moves at all, sitting near 0.99 throughout.
It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.18 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.75 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 63 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.809 before conversion and 0.866 after — it rose by 0.058. Neighbour-to-neighbour the worst pair went 0.701 → 0.872. (The earlier render, with segment 1 left raw, scores 0.756 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.348 in the original and +0.380 after conversion — 109 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Doubt, -0.033 became +0.150.
Quality. Mean predicted overall quality across the segments went 2.47 → 3.03 (+0.56) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.809 → 0.866+0.058identity cos neighbours 0.701 → 0.872d_b rescored +0.348 → +0.380d_a rescored -0.033 → +0.150d_a mined -0.036d_b mined 0.350min_cos_consec (site) 0.7496min_cos_anchor (site) 0.8271dataset podcastlang enspeaker 886160total 62.2schain gain +2.2 dBseam step 1.3 dBcrossfades 100/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a child strongly feminine voice · fairly smooth, below-average recording, fast, energised, neutral tension, moderately variable, some disfluency, very wide pitch range
(doubt, interest, intoxication altered states of consciousness · somewhat unclear, cartoonish, casual)in African uh (childlike giggle) in (ahem) uh in (ahem) uh Africa uh the half percentage of of the of (childlike giggle) uh Africa is Muslim. Yes. Yes, yes. Half percentage, but like uh (childlike giggle) Nigeria is filled with Muslims. I gotta say that, okay? And there's Morocco in Africa that's filled with Muslims, literally. They wear like Muslim dresses, they have a bunch of masjids. We (childlike giggle) uh so other places might have like that's pretty egoist.
full caption & clip details
A child strongly feminine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as doubt, interest, intoxication altered states of consciousness; style: cartoonish, casual; below-average recording, quiet background; genuineness 2.8/6; vocal-burst blend 6.6/10; 28.8s, EN.
886160_00250647 · in -18.7 dBFS · gain -1.3 dB · podcast-04282
(shame, longing, pain·slurred, casual, storytelling)But our God has told us not to change your skin the way that I sent you in. The way I made you is the way I made you and I don't want you to change your skin. Don't change yourself. And one And you know, it is actually haram. Haram means (ahem) um like
full caption & clip details
A child strongly feminine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is cool, bright, fairly smooth, slightly thin; slurred, some disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as shame, longing, pain; style: casual, storytelling; below-average recording, some background noise; genuineness 2.6/6; vocal-burst blend 3.5/10; 15.9s, EN.
886160_00253975 · in -21.0 dBFS · gain +1.0 dB · podcast-00489
(impatience and irritability, disappointment, jealousy and envy· slurred, casual, cartoonish)Yeah, but all I have to say just to wrap up this episode, make sure you guys like y you have to feel for the people. Because the people, like they they they're poor people right now. Thinking of like you have this mic right now. Like look at this mic. You this is all this is actually thirty dollars. And this
full caption & clip details
A child strongly feminine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is cool, slightly bright, fairly smooth, thin; slurred, some disfluency, very wide pitch range, normal breath; affect is elated, slightly dominant, neutral openness; reads as impatience and irritability, disappointment, jealousy and envy; style: casual, cartoonish; below-average recording, some background noise; genuineness 3.0/6; vocal-burst blend 5.7/10; 17.8s, EN.
886160_00257543 · in -16.8 dBFS · gain -3.2 dB · podcast-00485
This chain comes from the one-sided rule: only Embarrassment had to get where it was going, by at least 0.50. The other emotion was left completely free.
The chain starts with Embarrassment around average — 0.45, lower than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.54.
Nothing was asked of the other axis, and in fact Fatigue Exhaustion barely moves at all, sitting near 1.00 throughout.
It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.08, then +0.23 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.40 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.24 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.40, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 80 s · de · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.385 before conversion and 0.714 after — it rose by 0.329. Neighbour-to-neighbour the worst pair went 0.328 → 0.686. (The earlier render, with segment 1 left raw, scores 0.697 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.589 in the original and +0.622 after conversion — 106 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Fatigue Exhaustion, -0.043 became -0.040.
Quality. Mean predicted overall quality across the segments went 2.71 → 3.00 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.385 → 0.714+0.329identity cos neighbours 0.328 → 0.686d_b rescored +0.589 → +0.622d_a rescored -0.043 → -0.040d_a mined -0.047d_b mined 0.544min_cos_consec (site) 0.2446min_cos_anchor (site) 0.3976dataset podcastlang despeaker 675206total 78.8schain gain +1.0 dBseam step 2.2 dBcrossfades 150/150/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a middle-aged somewhat feminine voice · balanced body, quiet background, very low-energy, relaxed
(fatigue exhaustion, contentment, relief · slow, fairly steady, little disfluency, ASMR)Das schafft schon so ein paar Stunden Luft am Tag, wo man wirklich so in Ruhe arbeiten kann. Das Büro ist aber wirklich so drei Minuten von der Kita entfernt. Das heißt, man kann auch Kinder abholen und dann mit ins Büro nehmen. So und das passiert auch, das machen die auch. Und ich bin ja auch selbstständig. Das interessiert ja keinen, ob ich jetzt zwischen nachmittags um vier und abends um sechs was fertig schreibe. Oder ob ich (wistful sigh)
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; slurred, little disfluency, narrow pitch range, light breath; affect is mildly positive, submissive, neutral openness; reads as fatigue exhaustion, contentment, relief; style: ASMR, whispered; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 1.8/10; 28.8s, DE.
675206_00099096 · in -38.3 dBFS · gain +18.3 dB · podcast-01633
(jealousy and envy, affection, thankfulness gratitude· slow, fairly steady, frequent disfluency, ASMR)Und ansonsten haben wir natürlich irgendwie auch Babysitter oder ich habe auch relativ viel Familie. Also die Kinder fahren dann auch in den Sommerferien irgendwie. Zahl auch mit der Uma in Urlaub und so. Also, das sind schon alles Dinge, die natürlich auch passieren und helfen. Und das wäre, also es wäre nicht vorstellbar ohne Hilfe,
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, narrow pitch range, light breath; affect is mildly positive, submissive, neutral openness; reads as jealousy and envy, affection, thankfulness gratitude; style: ASMR, whispered; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 2.4/10; 23.0s, DE.
675206_00104611 · in -36.9 dBFS · gain +16.9 dB · podcast-01627
(contentment, doubt, affection ·normal-paced, moderately variable, frequent disfluency, casual)glaube ich, das zu machen. Du bist ja auch hier unter anderem, weil du mit Scheißgefühlen im weitesten Sinne auch schon Erfahrungen gemacht hast. Magst du mal erzählen, (ahem) wann du an so psychische Belastungsgrenzen gekommen bist und wie die sich bei dir äußern und was da so deine Geschichte dazu
full caption & clip details
A young adult somewhat feminine voice; delivery is very low-energy, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contentment, doubt, affection; style: casual, conversational; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 0.4/10; 21.5s, DE.
675206_00106928 · in -34.3 dBFS · gain +14.3 dB · podcast-03353
(embarrassment, shame, sadness·measured, moderately variable, frequent disfluency, casual)Also mir ist ja mal was sehr Lustiges aufgefallen. (wistful sigh) Ich glaube, ich war (ahem)
full caption & clip details
A young adult somewhat feminine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is mildly negative, submissive, neutral openness; reads as embarrassment, shame, sadness; style: casual, conversational; below-average recording, quiet background; genuineness 3.5/6; vocal-burst blend 0.0/10; 6.0s, DE.
675206_00109216 · in -32.3 dBFS · gain +12.3 dB · podcast-03346
This chain comes from the one-sided rule: only Bitterness had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Bitterness around average — 0.52, higher than 52 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.43.
Nothing was asked of the other axis, and in fact Relief drifts down from 0.99 to 0.32 (-0.66), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.05, then +0.24, then +0.14 — a plateau around step 1, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.35 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.36 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.35, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 59 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.414 before conversion and 0.769 after — it rose by 0.355. Neighbour-to-neighbour the worst pair went 0.297 → 0.663. (The earlier render, with segment 1 left raw, scores 0.611 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Bitterness moved +0.431 in the original and +0.612 after conversion — 142 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Relief, -0.661 became -0.187.
Quality. Mean predicted overall quality across the segments went 2.72 → 3.21 (+0.49) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.414 → 0.769+0.355identity cos neighbours 0.297 → 0.663d_b rescored +0.431 → +0.612d_a rescored -0.661 → -0.187d_a mined -0.662d_b mined 0.431min_cos_consec (site) 0.3556min_cos_anchor (site) 0.3479dataset podcastlang enspeaker 550148total 58.4schain gain +3.7 dBseam step 1.6 dBcrossfades 100/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body
(relief, pride, triumph · normal-paced, normally alert, relaxed, casual)Monday. At the time is the final (low mumble) month and the people with all the reset of the offer.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as relief, pride, triumph; style: casual, conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 4.3/10; 9.6s, EN.
550148_00182528 · in -30.5 dBFS · gain +10.5 dB · podcast-02745
(amusement, teasing, pleasure ecstasy· normal-paced, energised, neutral tension, casual)I've seen a meme that said when the interest of the 2022 are 18 years of interest. Exactly.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, no audible breath; affect is positive, slightly submissive, slightly guarded; reads as amusement, teasing, pleasure ecstasy; style: casual, playful; below-average recording, some background noise; genuineness 5.5/6; vocal-burst blend 5.3/10; 20.8s, EN.
550148_00183816 · in -29.7 dBFS · gain +9.7 dB · podcast-00374
(triumph, relief, elation· normal-paced, normally alert, neutral tension, conversational)And the ultimate thing is the thing of Black Friday, what do you think of the one thing? Then in Black Friday recuperates, we're at Cyber Monday. The people and now the Navidad, right? Exactly.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as triumph, relief, elation; style: conversational, casual; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 7.0/10; 23.7s, EN.
550148_00186080 · in -29.5 dBFS · gain +9.5 dB · podcast-06088
(bitterness, disgust·fast, normally alert, slightly relaxed, casual)hasta until the buy. There are other employees that pay.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as bitterness, disgust; style: casual, conversational; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 6.3/10; 4.8s, EN.
550148_00189400 · in -33.7 dBFS · gain +13.7 dB · podcast-01434
This chain comes from the one-sided rule: only Sexual Lust had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Sexual Lust clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.34.
Nothing was asked of the other axis, and in fact Amusement drifts down from 0.99 to 0.81 (-0.18), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.12, then +0.10, then +0.12, then -0.00 — not a clean run: step 4 moves back the other way by 0.00 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.39 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.56 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.39, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 48 s · en · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.268 before conversion and 0.580 after — it rose by 0.313. Neighbour-to-neighbour the worst pair went 0.541 → 0.754. (The earlier render, with segment 1 left raw, scores 0.492 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sexual Lust moved +0.343 in the original and +0.831 after conversion — 242 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Amusement, -0.186 became -0.071.
Quality. Mean predicted overall quality across the segments went 2.81 → 2.96 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.268 → 0.580+0.313identity cos neighbours 0.541 → 0.754d_b rescored +0.343 → +0.831d_a rescored -0.186 → -0.071d_a mined -0.184d_b mined 0.337min_cos_consec (site) 0.5586min_cos_anchor (site) 0.3931dataset podcastlang enspeaker 290495total 46.9schain gain +3.9 dBseam step 1.9 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, light breath
(amusement, intoxication altered states of consciousness, teasing · normal-paced, normally alert, relaxed, casual)coolest Batmobile. Man, I I really I I got the Lego set of the the Patents in Batman. I really like that 'cause it's basically like hey, what's the most badass like (low mumble) uh (low mumble) um oh what do you call those? Uh muscle car you can build and (low mumble) uh I I appreciate that one. But yeah, this one I think is like because this one is like wholly unique.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is neutral, slightly submissive, neutral openness; reads as amusement, intoxication altered states of consciousness, teasing; style: casual, playful; average recording, quiet background; genuineness 5.7/6; vocal-burst blend 4.5/10; 25.0s, EN.
290495_00283968 · in -33.6 dBFS · gain +13.6 dB · podcast-05970
(teasing, sourness, amusement · normal-paced, energised, neutral tension, casual)Like they didn't just like throw something on top of a chassis, they had to build a whole new vehicle. Or is this based on anything or they just designed this purely for this movie?
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as teasing, sourness, amusement; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.9/6; vocal-burst blend 3.5/10; 9.6s, EN.
290495_00286536 · in -33.9 dBFS · gain +13.9 dB · podcast-02663
(normal-paced, normally alert, slightly relaxed, casual)As far as I know, it was purely for the movie because I know that Jay Leno had it. Okay,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; explicit content; genuineness 4.0/6; vocal-burst blend 1.9/10; 4.6s, EN.
290495_00287520 · in -37.7 dBFS · gain +17.7 dB · podcast-02657
(sexual lust· normal-paced, normally alert, fully relaxed, casual)it's just has a V eight like small block engine in it, but it
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as sexual lust; style: casual, conversational; average recording, quiet background; explicit content; genuineness 4.8/6; vocal-burst blend 1.4/10; 3.7s, EN.
290495_00288248 · in -39.5 dBFS · gain +19.5 dB · podcast-02657
(sexual lust, disgust, intoxication altered states of consciousness·slow, normally alert, slightly relaxed, casual)has an psychotically low gear ratio so it can turn the tires.
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as sexual lust, disgust, intoxication altered states of consciousness; style: casual, whispered; average recording, no background noise; explicit content; genuineness 2.8/6; vocal-burst blend 2.4/10; 4.8s, EN.
290495_00288616 · in -40.3 dBFS · gain +20.3 dB · podcast-02688
This chain comes from the one-sided rule: only Confusion had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Confusion strongly present — 0.77, higher than 77 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.22.
Nothing was asked of the other axis, and in fact Contemplation barely moves at all, sitting near 0.99 throughout.
It takes 2 clips to get there. Clip to clip the moves are +0.22 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.02 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.02 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.02, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 15 s · en · podcast
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.050 before conversion and 0.725 after — it rose by 0.675. Neighbour-to-neighbour the worst pair went 0.050 → 0.725. (The earlier render, with segment 1 left raw, scores 0.484 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.216 in the original and +0.349 after conversion — 162 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contemplation, -0.048 became -0.045.
Quality. Mean predicted overall quality across the segments went 2.80 → 3.00 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.050 → 0.725+0.675identity cos neighbours 0.050 → 0.725d_b rescored +0.216 → +0.349d_a rescored -0.048 → -0.045d_a mined -0.048d_b mined 0.222min_cos_consec (site) 0.0219min_cos_anchor (site) 0.0219dataset podcastlang enspeaker 585382total 15.0schain gain +4.1 dBseam step 1.0 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, moderate pitch range
(contemplation, fear, distress · slightly relaxed, fairly steady, some disfluency, casual)And 'cause I think that's what scares me most about death is the idea of being completely
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as contemplation, fear, distress; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 1.5/10; 5.5s, EN.
585382_00324192 · in -29.5 dBFS · gain +9.5 dB · podcast-04317
(confusion, pain, shame·relaxed, moderately variable, frequent disfluency, casual)and she said, you know, we don't hear that often, we don't hear how we're gonna be remembered ever. You know, and that's what that video was. And like I thought
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as confusion, pain, shame; style: casual, conversational; good recording, quiet background; genuineness 3.7/6; vocal-burst blend 6.5/10; 9.7s, EN.
585382_00328832 · in -29.6 dBFS · gain +9.6 dB · podcast-04334
This chain comes from the one-sided rule: only Longing had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Longing clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.31.
Nothing was asked of the other axis, and in fact Concentration drifts down from 0.98 to 0.52 (-0.46), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.11, then +0.22, then -0.01 — not a clean run: step 3 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.10 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.17 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.10, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 45 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.101 before conversion and 0.677 after — it rose by 0.576. Neighbour-to-neighbour the worst pair went 0.160 → 0.531. (The earlier render, with segment 1 left raw, scores 0.445 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.314 in the original and +0.136 after conversion — 44 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Concentration, -0.462 became -0.321.
Quality. Mean predicted overall quality across the segments went 2.83 → 3.08 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.101 → 0.677+0.576identity cos neighbours 0.160 → 0.531d_b rescored +0.314 → +0.136d_a rescored -0.462 → -0.321d_a mined -0.464d_b mined 0.313min_cos_consec (site) 0.1725min_cos_anchor (site) 0.1017dataset podcastlang enspeaker 269604total 43.9schain gain +3.6 dBseam step 1.9 dBcrossfades 100/100/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, fairly steady, average clarity
(concentration, contemplation · subdued, relaxed, frequent disfluency, casual)place. So and I do want to add on that too. So, like my relationship with my mentor is it's not like, hey, I have a question, can you give me the answer? It's more of like I present the question. This is what I'm thinking about doing. What are your thoughts? So
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as concentration, contemplation; style: casual, conversational; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 9.9/10; 15.2s, EN.
269604_00150112 · in -29.1 dBFS · gain +9.2 dB · podcast-04547
(fatigue exhaustion, emotional numbness, contemplation ·normally alert, slightly relaxed, some disfluency, casual)getting like a second opinion or more or less co-authoring the solution so that you're not just like, give me the answer.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as fatigue exhaustion, emotional numbness, contemplation; style: casual, whispered; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 3.8/10; 7.6s, EN.
269604_00151752 · in -27.5 dBFS · gain +7.5 dB · podcast-04560
(contemplation, interest, longing· normally alert, slightly relaxed, some disfluency, casual)Yeah. And really, it's like the the relationship that I've had with him. And every time I like bring up a question, he he answers it with another question. And then at in the end, I knew the answer to the question that I was gonna ask.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contemplation, interest, longing; style: casual, conversational; good recording, quiet background; genuineness 4.9/6; vocal-burst blend 10.0/10; 13.3s, EN.
269604_00152688 · in -27.2 dBFS · gain +7.2 dB · podcast-04550
(longing, contemplation · normally alert, neutral tension, some disfluency, authoritative)(ahem) um, I've always liked to recap the previous meeting in a sense, letting them know that hey, I listened to the advice you gave me.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as longing, contemplation; style: authoritative, conversational; good recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.8/10; 8.3s, EN.
269604_00159332 · in -30.6 dBFS · gain +10.6 dB · podcast-04544
This chain comes from the one-sided rule: only Pride had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Pride clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.33.
Nothing was asked of the other axis, and in fact Amusement drifts down from 0.96 to 0.76 (-0.20), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.23, then +0.10 — a plateau around step 1, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.90 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 42 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.888 before conversion and 0.885 after — it fell by 0.003. Neighbour-to-neighbour the worst pair went 0.906 → 0.900. (The earlier render, with segment 1 left raw, scores 0.787 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.331 in the original and +0.249 after conversion — 75 % of the delta retained, which is most of it. On the other named axis, Amusement, -0.201 became -0.059.
Quality. Mean predicted overall quality across the segments went 2.82 → 2.97 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.888 → 0.885-0.003identity cos neighbours 0.906 → 0.900d_b rescored +0.331 → +0.249d_a rescored -0.201 → -0.059d_a mined -0.200d_b mined 0.328min_cos_consec (site) 0.8982min_cos_anchor (site) 0.8990dataset podcastlang enspeaker 148621total 41.5schain gain +4.6 dBseam step 1.3 dBcrossfades 150/100/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a child feminine voice · slightly bright, fairly smooth, quiet background, brisk, energised, neutral tension, moderately variable, some disfluency
(amusement, hope enthusiasm optimism, disgust · clear, normal breath, dramatic, cartoonish)The community thrived because everybody knew everybody couldn't cook. Everybody can't play the banjo. Everybody can't drive.
full caption & clip details
A child feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as amusement, hope enthusiasm optimism, disgust; style: dramatic, cartoonish; good recording, quiet background; genuineness 2.1/6; vocal-burst blend 2.0/10; 7.8s, EN.
148621_00325512 · in -22.7 dBFS · gain +2.7 dB · podcast-05961
(helplessness, sadness, fatigue exhaustion·average clarity, light breath, conversational, casual)Everybody had a task. But now nowadays I feel like everybody is losing the fact that everybody's an individual.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as helplessness, sadness, fatigue exhaustion; style: conversational, casual; good recording, quiet background; mildly explicit content; genuineness 4.1/6; vocal-burst blend 3.8/10; 7.8s, EN.
148621_00326464 · in -24.9 dBFS · gain +4.9 dB · podcast-05956
(jealousy and envy, impatience and irritability, sourness· average clarity, light breath, conversational, casual)And when you try to get everybody to be the same person and do the same thing, it may work for a couple of people, but it's always gonna be those people that break the mold.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as jealousy and envy, impatience and irritability, sourness; style: conversational, casual; average recording, quiet background; mildly explicit content; genuineness 3.0/6; vocal-burst blend 3.8/10; 10.2s, EN.
148621_00327680 · in -27.3 dBFS · gain +7.3 dB · podcast-05966
(pride, triumph, fear· average clarity, normal breath, casual, ranting)I have always known I was different. Do not try to put me with these people because it's not gonna work. I'm not gonna see it. It's not many people walking this earth that's gonna see stuff how I see it. And when I find those mirror, mirror, mirror, mirror people, they be around and they stay around. (ahem)
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, slightly guarded; reads as pride, triumph, fear; style: casual, ranting; average recording, quiet background; mildly explicit content; genuineness 4.4/6; vocal-burst blend 5.3/10; 16.2s, EN.
148621_00328696 · in -25.9 dBFS · gain +5.9 dB · podcast-05968
This chain comes from the one-sided rule: only Fatigue Exhaustion had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Fatigue Exhaustion clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.25.
Nothing was asked of the other axis, and in fact Teasing drifts down from 0.96 to 0.36 (-0.59), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.09 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 40 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.759 before conversion and 0.785 after — it rose by 0.026. Neighbour-to-neighbour the worst pair went 0.759 → 0.785. (The earlier render, with segment 1 left raw, scores 0.566 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.250 in the original and +0.232 after conversion — 93 % of the delta retained, which is essentially all of it. On the other named axis, Teasing, -0.596 became -0.154.
Quality. Mean predicted overall quality across the segments went 2.60 → 3.17 (+0.57) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.759 → 0.785+0.026identity cos neighbours 0.759 → 0.785d_b rescored +0.250 → +0.232d_a rescored -0.596 → -0.154d_a mined -0.594d_b mined 0.250min_cos_consec (site) 0.8315min_cos_anchor (site) 0.8418dataset podcastlang enspeaker 137262total 39.2schain gain +4.4 dBseam step 0.8 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · slightly cool, neutral-bright, slightly rough, normal-paced, neutral tension
(embarrassment, teasing · normally alert, moderately variable, some disfluency, casual)Because First of all, I wanna I want to have a discussion about my favorite wrestler right now. And that's Cody Rhodes.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; slurred, some disfluency, wide pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as embarrassment, teasing; style: casual, conversational; below-average recording, some background noise; genuineness 4.3/6; vocal-burst blend 2.0/10; 7.2s, EN.
137262_00217192 · in -34.3 dBFS · gain +14.3 dB · podcast-00771
(interest, elation, hope enthusiasm optimism·very low-energy, moderately variable, frequent disfluency, casual)They they go jump, they go sky with the Golden Knights. It's Alberto Del Rio, Coffee Kingston, CM Punk, Cody Rhodes. Uh (exhausted groan) come to find out, and this is what I love about heels. The two hills, the guys who are heels, are the two nicest guys in real life. Like Alberto Dorio's like actually really sweet and funny, and Cody Rhodes is just like a big kid. He like laughs, cuts up, and I'm just like it's so cool to see that.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as interest, elation, hope enthusiasm optimism; style: casual, playful; below-average recording, noisy background; mildly explicit content; genuineness 5.8/6; vocal-burst blend 9.7/10; 24.1s, EN.
137262_00219488 · in -34.3 dBFS · gain +14.3 dB · podcast-01760
(fatigue exhaustion, impatience and irritability, anger·energised, fairly steady, some disfluency, casual)Make a wish stuff, constantly back and forth, and like they work like two hundred and something days out of the year, and it's just like god of mine, unless you like have an injury.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as fatigue exhaustion, impatience and irritability, anger; style: casual, monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 4.0/10; 8.3s, EN.
137262_00224808 · in -36.8 dBFS · gain +16.8 dB · podcast-00761
This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Concentration clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.30.
Nothing was asked of the other axis, and in fact Pride drifts down from 0.99 to 0.00 (-0.99), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.09, then +0.15, then +0.05 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.77 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.76 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.77, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 102 s · de · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.817 before conversion and 0.873 after — it rose by 0.056. Neighbour-to-neighbour the worst pair went 0.823 → 0.879. (The earlier render, with segment 1 left raw, scores 0.699 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.282 in the original and +0.311 after conversion — 110 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Pride, -0.988 became -0.991.
Quality. Mean predicted overall quality across the segments went 2.93 → 3.14 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.817 → 0.873+0.056identity cos neighbours 0.823 → 0.879d_b rescored +0.282 → +0.311d_a rescored -0.988 → -0.991d_a mined -0.987d_b mined 0.297min_cos_consec (site) 0.7649min_cos_anchor (site) 0.7703dataset podcastlang despeaker 134525total 101.1schain gain +3.0 dBseam step 1.9 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(pride, thankfulness gratitude, contentment · casual, conversational)(breathy giggle) (ahem) Die Ausbildung habe ich tatsächlich dann in der Coaching-Akademie gemacht. And one of my absolute key Elemente, die ich auch in mein Unternehmen mit reinbringe, wo auch meine Assistentin und ich erst vor kurzem ein außerordentliches Telefonat geführt hatten, und ich das einem Klienten für seine beiden geschäftsführenden Partner
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as pride, thankfulness gratitude, contentment; style: casual, conversational; good recording, quiet background; genuineness 3.9/6; vocal-burst blend 0.4/10; 22.3s, DE.
134525_00264640 · in -21.9 dBFS · gain +1.9 dB · podcast-03547
(bitterness, disappointment, sadness·monologue, conversational)mitgegeben habe, ist eine gemeinsame Feedback-Kultur zu etablieren. Wenn es nicht am Anfang passiert ist, wenn es nicht zwischen Frau, Frau, Mann, Mann, Frau, Mann, Mann, Frau passiert ist, das darf zwischen allen passieren. Das darf in Familien passieren, zwischen Geschwistern, zwischen nur Geschäftspartnern für sich, zwischen allen hierarchischen Ebenen eines Unternehmens,
full caption & clip details
An adult somewhat feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as bitterness, disappointment, sadness; style: monologue, conversational; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 0.0/10; 26.3s, DE.
134525_00266866 · in -21.1 dBFS · gain +1.1 dB · podcast-03553
(sourness, hope enthusiasm optimism, contemplation· monologue, conversational)dass sich die beteiligten Parteien wirklich zusammenfinden und zusammen eine Feedback-Kultur etablieren. Und mit Feedback meine ich nicht nur, dass Feedback an sich, das hast du gut gemacht, dass nicht. Nein, nicht richtig falsch, das gibt es sowieso nicht. Auch nicht nur Rückmeldungen und keine Rückmeldung, sondern im Feedback ist für mich auch drin, in welchem Rahmen bewegen wir uns, welche Rahmen und Bedingungen hat unsere Zusammenarbeit, unsere Kommunikation in der Zusammenarbeit
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as sourness, hope enthusiasm optimism, contemplation; style: monologue, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 1.3/10; 29.4s, DE.
134525_00269493 · in -20.1 dBFS · gain +0.1 dB · podcast-02016
(concentration, disappointment, sadness· monologue, narration)und innerhalb dieser Rahmenbedingungen eben zu sagen, wem was wichtig ist. Und die Schlüsselelemente, die tragen, die werden dann zusammengetragen und wirklich immer wieder daran erinnert, groß aufgeschrieben oder von mir aus eingeblendet an der Stelle, wo morgens jeder vorbeigeht, so dass das Unbewusste das immer wieder mit aufnehmen kann.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, disappointment, sadness; style: monologue, narration; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 0.0/10; 23.7s, DE.
134525_00272427 · in -19.4 dBFS · gain -0.6 dB · podcast-02014
This chain comes from the one-sided rule: only Confusion had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Confusion clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.27.
Nothing was asked of the other axis, and in fact Embarrassment drifts down from 0.94 to 0.75 (-0.19), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.14, then +0.13 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.73 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.79 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.73, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 21 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.772 before conversion and 0.750 after — it fell by 0.022. Neighbour-to-neighbour the worst pair went 0.772 → 0.750. (The earlier render, with segment 1 left raw, scores 0.651 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.269 in the original and +0.114 after conversion — 42 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Embarrassment, -0.192 became -0.109.
Quality. Mean predicted overall quality across the segments went 2.66 → 3.00 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.772 → 0.750-0.022identity cos neighbours 0.772 → 0.750d_b rescored +0.269 → +0.114d_a rescored -0.192 → -0.109d_a mined -0.191d_b mined 0.269min_cos_consec (site) 0.7854min_cos_anchor (site) 0.7252dataset podcastlang enspeaker 863187total 20.1schain gain +2.9 dBseam step 1.5 dBcrossfades 100/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, average recording, quiet background, normally alert
(embarrassment · measured, relaxed, moderately variable, casual)No, no. No. I I have I guess anecdotal evidence for this. Or like my
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, relaxed, moderately variable; timbre is neutral-toned, dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, submissive, neutral openness; reads as embarrassment; style: casual, conversational; average recording, quiet background; genuineness 5.8/6; vocal-burst blend 0.8/10; 5.2s, EN.
863187_00490728 · in -23.1 dBFS · gain +3.1 dB · podcast-04225
(embarrassment, intoxication altered states of consciousness·normal-paced, neutral tension, moderately variable, casual)(low mumble) Uh well, you know, a personal experience due to this. So I have two dogs. Uh (low mumble) as Andrew knows. (low mumble) Uh I have my small dog, uh (low mumble) his name is Jack. And no
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; reads as embarrassment, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 4.3/10; 9.8s, EN.
863187_00491392 · in -25.0 dBFS · gain +5.0 dB · podcast-01398
(confusion, intoxication altered states of consciousness, disgust· normal-paced, relaxed, fairly steady, casual)his name is George. Small dog's name is George. He's like a (ahem) chihuahua mixed with I don't know what.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as confusion, intoxication altered states of consciousness, disgust; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.8/6; vocal-burst blend 5.5/10; 5.4s, EN.
863187_00492368 · in -24.9 dBFS · gain +4.9 dB · podcast-01382