c-evasnippets-B1 — voice-corrected

Corpus evasnippets in isolation, rule B1.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_c-evasnippets-B1.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
59segments re-voiced
0.769 → 0.846median worst-to-anchor identity cosine
95 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Relief(unconstrained axis: Disappointment)identity +0.05 emotion 72 %   c-evasnippets-B1 · #1

This chain comes from the one-sided rule: only Relief had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Relief strongly present — 0.76, higher than 76 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Disappointment drifts down from 0.98 to 0.87 (-0.11), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.00, then +0.00 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 78 s · en · evasnippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.850 before conversion and 0.895 after — it rose by 0.045. Neighbour-to-neighbour the worst pair went 0.850 → 0.895. (The earlier render, with segment 1 left raw, scores 0.722 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.240 in the original and +0.173 after conversion — 72 % of the delta retained, which is most of it. On the other named axis, Disappointment, -0.112 became -0.066.

Quality. Mean predicted overall quality across the segments went 2.99 → 3.13 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.850 → 0.895 +0.045identity cos neighbours 0.850 → 0.895d_b rescored +0.240 → +0.173d_a rescored -0.112 → -0.066d_a mined -0.112d_b mined 0.240min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_00_01_0002_5total 76.8schain gain +6.0 dBseam step 0.3 dBcrossfades 100/100/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, slightly bright, fairly smooth, balanced body, moderately variable, some disfluency, average clarity, light breath
(disappointment, bitterness, affection · normal-paced, normally alert, slightly relaxed, monologue) It's always the same because like I always say, kids are kids, our kids are kids. Wherever you go, kids are the same. They want to be seen, they want to be loved, they want to be heard, they want to be valued, they want to be acknowledged, they want to be validated. They want to know that they're loved unconditionally, not based on anything they do or don't do. So that's why I'm always trying to have you as parents not look at their grades, not look at their ADD, not look at their sports.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as disappointment, bitterness, affection; style: monologue, casual; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 5.5/10; 27.0s.
cond_podcastt_00_01_0002_518189_00011016 · in -19.3 dBFS · gain -0.7 dB · evasnippets-00324
(relief, amusement, fear · brisk, energised, neutral tension, casual) I would just hear the water running in the next room and I would shoot out of bed because he did give us a couple warnings. And then finally he would just come and just drip, drip, drip. And you're like, oh, why did I wait? Why did I wait? So then the next time there was no waiting. Tell them about the time when I was trying to get you back with the water.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as relief, amusement, fear; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 4.3/6; vocal-burst blend 8.1/10; 16.7s.
cond_podcastt_00_01_0002_518189_00046448 · in -19.0 dBFS · gain -1.0 dB · evasnippets-00316
(relief, amusement, fear · brisk, energised, neutral tension, casual) I would just hear the water running in the next room and I would shoot out of bed because he did give us a couple warnings. And then finally he would just come and just drip, drip, drip. And you're like, oh, why did I wait? Why did I wait? So then the next time there was no waiting. Tell them about the time when I was trying to get you back with the water.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as relief, amusement, fear; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 4.3/6; vocal-burst blend 8.1/10; 16.7s.
cond_podcastt_00_01_0002_518189_00046448 · in -19.0 dBFS · gain -1.0 dB · evasnippets-00316
(relief, amusement, fear · brisk, energised, neutral tension, casual) I would just hear the water running in the next room and I would shoot out of bed because he did give us a couple warnings. And then finally he would just come and just drip, drip, drip. And you're like, oh, why did I wait? Why did I wait? So then the next time there was no waiting. Tell them about the time when I was trying to get you back with the water.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as relief, amusement, fear; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 4.3/6; vocal-burst blend 8.1/10; 16.7s.
cond_podcastt_00_01_0002_518189_00046448 · in -19.0 dBFS · gain -1.0 dB · evasnippets-00316
Pride(unconstrained axis: Disappointment)identity −0.01 emotion 102 %   c-evasnippets-B1 · #2

This chain comes from the one-sided rule: only Pride had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Pride clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Disappointment drifts down from 1.00 to 0.39 (-0.60), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.00, then +0.00 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 87 s · en · evasnippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.851 before conversion and 0.843 after — it fell by 0.009. Neighbour-to-neighbour the worst pair went 0.851 → 0.843. (The earlier render, with segment 1 left raw, scores 0.662 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.239 in the original and +0.243 after conversion — 102 % of the delta retained, which is essentially all of it. On the other named axis, Disappointment, -0.601 became -0.595.

Quality. Mean predicted overall quality across the segments went 3.29 → 3.45 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.851 → 0.843 -0.009identity cos neighbours 0.851 → 0.843d_b rescored +0.239 → +0.243d_a rescored -0.601 → -0.595d_a mined -0.601d_b mined 0.239min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_00_00_0000_2total 86.2schain gain +2.4 dBseam step 0.4 dBcrossfades 150/150/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, measured, subdued, slightly relaxed
(disappointment, helplessness, sadness · almost no disfluency, narration, monologue) They they don't want to hear about it anymore and they they'll have their stock kind of comebacks. Like your kids are never gonna be able to socialize, they'll never be able to get ahead in life, they're never gonna get a job, they'll never be able to go to university. And all of this programming stuff that you see, and it's very difficult to to face that.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; slurred, almost no disfluency, moderate pitch range, minimal breath; affect is mildly negative, neutral stance, slightly guarded; reads as disappointment, helplessness, sadness; style: narration, monologue; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 4.7/10; 20.4s.
cond_podcastt_00_00_0000_290845_00320772 · in -33.6 dBFS · gain +13.6 dB · evasnippets-00324
(pride, doubt, triumph · some disfluency, narration, monologue) who uh (low mumble) Scott, who was supposed to obviously be very good academic as well, (low mumble) um, basically got him s himself expelled from school at the age of eleven and Peter couldn't figure out why, you know, what's going on. So he fell down the rabbit hole of self directed education in democratic schooling and uh (low mumble) found the Sudbury Valley School, uh, (ahem) which is like uh (surprised gasp) the the probably the most well known democratic school in
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, minimal breath; affect is mildly negative, neutral stance, slightly guarded; reads as pride, doubt, triumph; style: narration, monologue; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 5.3/10; 22.1s.
cond_podcastt_00_00_0000_290845_00371460 · in -33.8 dBFS · gain +13.8 dB · evasnippets-00324
(pride, doubt, triumph · some disfluency, narration, monologue) who uh (low mumble) Scott, who was supposed to obviously be very good academic as well, (low mumble) um, basically got him s himself expelled from school at the age of eleven and Peter couldn't figure out why, you know, what's going on. So he fell down the rabbit hole of self directed education in democratic schooling and uh (low mumble) found the Sudbury Valley School, uh, (ahem) which is like uh (surprised gasp) the the probably the most well known democratic school in
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, minimal breath; affect is mildly negative, neutral stance, slightly guarded; reads as pride, doubt, triumph; style: narration, monologue; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 5.3/10; 22.1s.
cond_podcastt_00_00_0000_290845_00371460 · in -33.8 dBFS · gain +13.8 dB · evasnippets-00324
(pride, doubt, triumph · some disfluency, narration, monologue) who uh (low mumble) Scott, who was supposed to obviously be very good academic as well, (low mumble) um, basically got him s himself expelled from school at the age of eleven and Peter couldn't figure out why, you know, what's going on. So he fell down the rabbit hole of self directed education in democratic schooling and uh (low mumble) found the Sudbury Valley School, uh, (ahem) which is like uh (surprised gasp) the the probably the most well known democratic school in
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, minimal breath; affect is mildly negative, neutral stance, slightly guarded; reads as pride, doubt, triumph; style: narration, monologue; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 5.3/10; 22.1s.
cond_podcastt_00_00_0000_290845_00371460 · in -33.8 dBFS · gain +13.8 dB · evasnippets-00324
Teasing(unconstrained axis: Infatuation)identity −0.01 emotion 99 %   c-evasnippets-B1 · #3

This chain comes from the one-sided rule: only Teasing had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Teasing strongly present — 0.80, higher than 80 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.20.

Nothing was asked of the other axis, and in fact Infatuation barely moves at all, sitting near 1.00 throughout.

It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.20, then +0.00 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 99 s · en · evasnippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.781 before conversion and 0.769 after — it fell by 0.011. Neighbour-to-neighbour the worst pair went 0.781 → 0.789. (The earlier render, with segment 1 left raw, scores 0.697 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Teasing moved +0.200 in the original and +0.199 after conversion — 99 % of the delta retained, which is essentially all of it. On the other named axis, Infatuation, +0.001 became -0.002.

Quality. Mean predicted overall quality across the segments went 2.96 → 3.17 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.781 → 0.769 -0.011identity cos neighbours 0.781 → 0.789d_b rescored +0.200 → +0.199d_a rescored +0.001 → -0.002d_a mined 0.001d_b mined 0.200min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_00_01_0003_6total 97.6schain gain +5.4 dBseam step 0.9 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a middle-aged somewhat feminine voice · slightly cool, neutral-bright, slightly rough, average recording, quiet background, normally alert
(infatuation, disgust, awe · slow, slightly relaxed, fairly steady, whispered) crammed beneath the visor of a cap. The face consisted of a rapid nose, droopy moustache, ferocious watery small eyes, a pugnacious chin and sunken cheeks, hideously smiling. There was something in the ensemble, at once brutal and ridiculous, vigorous and pathetic.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, thin; clear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, neutral stance, slightly guarded; reads as infatuation, disgust, awe; style: whispered, narration; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 0.1/10; 22.8s.
cond_podcastt_00_01_0003_626648_00078540 · in -31.4 dBFS · gain +11.4 dB · evasnippets-00322
(infatuation, disgust, awe · slow, slightly relaxed, fairly steady, whispered) crammed beneath the visor of a cap. The face consisted of a rapid nose, droopy moustache, ferocious watery small eyes, a pugnacious chin and sunken cheeks, hideously smiling. There was something in the ensemble, at once brutal and ridiculous, vigorous and pathetic.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, thin; clear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, neutral stance, slightly guarded; reads as infatuation, disgust, awe; style: whispered, narration; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 0.1/10; 22.8s.
cond_podcastt_00_01_0003_626648_00078540 · in -31.4 dBFS · gain +11.4 dB · evasnippets-00322
(teasing, pleasure ecstasy, infatuation · brisk, neutral tension, moderately variable, casual) I say I do your sweep for you, it translated pleasantly. I thanked it, and the vulture, exclaiming, Good, good not me, Savrillon. Haret does it for everybody (childlike giggle) rushed off, followed by Haret and the tassel. Out of the corner of my eye I watched the tall, ludicrous, extraordinary, almost proud figure of the bear stoop with quiet dignity.
full caption & clip details
An elderly somewhat feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as teasing, pleasure ecstasy, infatuation; style: casual, cartoonish; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.4/10; 26.3s.
cond_podcastt_00_01_0003_626648_00083520 · in -27.6 dBFS · gain +7.6 dB · evasnippets-00322
(teasing, pleasure ecstasy, infatuation · brisk, neutral tension, moderately variable, casual) I say I do your sweep for you, it translated pleasantly. I thanked it, and the vulture, exclaiming, Good, good not me, Savrillon. Haret does it for everybody (childlike giggle) rushed off, followed by Haret and the tassel. Out of the corner of my eye I watched the tall, ludicrous, extraordinary, almost proud figure of the bear stoop with quiet dignity.
full caption & clip details
An elderly somewhat feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as teasing, pleasure ecstasy, infatuation; style: casual, cartoonish; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.4/10; 26.3s.
cond_podcastt_00_01_0003_626648_00083520 · in -27.6 dBFS · gain +7.6 dB · evasnippets-00322
Confusion(unconstrained axis: Sexual Lust)identity +0.04 emotion 34 %   c-evasnippets-B1 · #4

This chain comes from the one-sided rule: only Confusion had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Confusion clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.25.

Nothing was asked of the other axis, and in fact Sexual Lust drifts down from 0.99 to 0.78 (-0.21), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 29 s · en · evasnippets

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.885 before conversion and 0.927 after — it rose by 0.042. Neighbour-to-neighbour the worst pair went 0.885 → 0.927. (The earlier render, with segment 1 left raw, scores 0.779 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.246 in the original and +0.083 after conversion — 34 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Sexual Lust, -0.212 became -0.171.

Quality. Mean predicted overall quality across the segments went 2.91 → 3.20 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.885 → 0.927 +0.042identity cos neighbours 0.885 → 0.927d_b rescored +0.246 → +0.083d_a rescored -0.212 → -0.171d_a mined -0.212d_b mined 0.247min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_03_02_0000_8total 28.3schain gain +2.8 dBseam step 0.4 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, slightly bright, fairly smooth, balanced body, good recording, quiet background, normal-paced, neutral tension
(sexual lust, pleasure ecstasy, jealousy and envy · energised, some disfluency, casual, playful) You know, I was on the team, so my it's my duty to also do the same thing. So I was scouring through all the photos of like us playing football. And I found this one photo. Maybe I can find it.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as sexual lust, pleasure ecstasy, jealousy and envy; style: casual, playful; good recording, quiet background; genuineness 6.0/6; vocal-burst blend 5.3/10; 11.9s.
cond_podcastt_03_02_0000_866755_00054536 · in -39.0 dBFS · gain +19.0 dB · evasnippets-00299
(confusion, embarrassment, amusement · normally alert, frequent disfluency, casual, playful) bonded together. I think I don't know exactly how the conversation happened, so you'll have to remind me or remind the the listeners. (contented sigh) Um but we um (ahem) I was looking for a job and you just said like hey there's a job here and is that how that happened? Is that how that works?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as confusion, embarrassment, amusement; style: casual, playful; good recording, quiet background; genuineness 5.5/6; vocal-burst blend 6.3/10; 16.6s.
cond_podcastt_03_02_0000_866755_00068320 · in -35.3 dBFS · gain +15.3 dB · evasnippets-00072
Thankfulness Gratitude(unconstrained axis: Interest)identity +0.64 emotion 76 %   c-evasnippets-B1 · #5

This chain comes from the one-sided rule: only Thankfulness Gratitude had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Thankfulness Gratitude strongly present — 0.78, higher than 78 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, virtually no clip in this corpus scores higher. That is a total rise of 0.21.

Nothing was asked of the other axis, and in fact Interest drifts down from 0.97 to 0.75 (-0.22), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are -0.01, then +0.22, then +0.00, then +0.00 — not a clean run: step 1 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 132 s · en · evasnippets

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.103 before conversion and 0.745 after — it rose by 0.643. Neighbour-to-neighbour the worst pair went 0.110 → 0.812. (The earlier render, with segment 1 left raw, scores 0.598 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.212 in the original and +0.162 after conversion — 76 % of the delta retained, which is most of it. On the other named axis, Interest, -0.216 became -0.135.

Quality. Mean predicted overall quality across the segments went 3.07 → 3.40 (+0.32) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.103 → 0.745 +0.643identity cos neighbours 0.110 → 0.812d_b rescored +0.212 → +0.162d_a rescored -0.216 → -0.135d_a mined -0.215d_b mined 0.212min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_00_03_0000_7total 130.7schain gain +1.4 dBseam step 0.8 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, slightly rough, balanced body, average recording, quiet background, subdued, slightly relaxed
(interest, bitterness · measured, some disfluency, somewhat unclear, casual) In the uh (low mumble) modern English is supervisor or manager, you know. Um, (low mumble) but overseer has been translated bishop and turned into a hierarchical thing in most parts of the world. In actual fact, it says the elders, you are overseers. So they're what they're overseers. Now, here's the key point: an overseer doesn't do the work, they watch people do work their work.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as interest, bitterness; style: casual, monologue; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 5.4/10; 21.6s.
cond_podcastt_00_03_0000_776279_00239856 · in -27.1 dBFS · gain +7.1 dB · evasnippets-00154
(concentration, bitterness, shame · measured, frequent disfluency, average clarity, monologue) (low mumble) uh plants in a person something that doesn't fit the mold. As that little believer tries to bring it to fulfillment, those strong through uh (ahem) those strong through human religion and tradition try to kill it. Too often they succeed. We need to humbly, meekly, blindly follow the Holy Spirit in such a way that the traditions in us that are offensive to God will be torn down. Man, my brother
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, bitterness, shame; style: monologue, casual; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.1/10; 26.0s.
cond_podcastt_00_03_0000_776279_00277816 · in -27.2 dBFS · gain +7.2 dB · evasnippets-00012
(thankfulness gratitude, contentment, affection · normal-paced, some disfluency, somewhat unclear, casual) Like we cannot imagine. Amen. Amen. And I will just add to all of that that as you launch out, folks, just remain accountable and uh (low mumble) come alongside another person who can mentor you and help you and guide you. You know, this book, David, as you mentioned to me, really also grew out of those who mentored you and poured into your life and helped you to understand these deep concepts. And I love that, brother. So thank you for writing this book. Thank you for being obedient.
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as thankfulness gratitude, contentment, affection; style: casual, monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 10.0/10; 28.0s.
cond_podcastt_00_03_0000_776279_00307112 · in -26.7 dBFS · gain +6.7 dB · evasnippets-00326
(thankfulness gratitude, contentment, affection · normal-paced, some disfluency, somewhat unclear, casual) Like we cannot imagine. Amen. Amen. And I will just add to all of that that as you launch out, folks, just remain accountable and uh (low mumble) come alongside another person who can mentor you and help you and guide you. You know, this book, David, as you mentioned to me, really also grew out of those who mentored you and poured into your life and helped you to understand these deep concepts. And I love that, brother. So thank you for writing this book. Thank you for being obedient.
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as thankfulness gratitude, contentment, affection; style: casual, monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 10.0/10; 28.0s.
cond_podcastt_00_03_0000_776279_00307112 · in -26.7 dBFS · gain +6.7 dB · evasnippets-00326
(thankfulness gratitude, contentment, affection · normal-paced, some disfluency, somewhat unclear, casual) Like we cannot imagine. Amen. Amen. And I will just add to all of that that as you launch out, folks, just remain accountable and uh (low mumble) come alongside another person who can mentor you and help you and guide you. You know, this book, David, as you mentioned to me, really also grew out of those who mentored you and poured into your life and helped you to understand these deep concepts. And I love that, brother. So thank you for writing this book. Thank you for being obedient.
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as thankfulness gratitude, contentment, affection; style: casual, monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 10.0/10; 28.0s.
cond_podcastt_00_03_0000_776279_00307112 · in -26.7 dBFS · gain +6.7 dB · evasnippets-00326
Interest(unconstrained axis: Affection)identity +0.66 emotion 113 %   c-evasnippets-B1 · #6

This chain comes from the one-sided rule: only Interest had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Interest strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.20.

Nothing was asked of the other axis, and in fact Affection barely moves at all, sitting near 1.00 throughout.

It takes 5 clips to get there. Clip to clip the moves are +0.00, then +0.20, then +0.00, then +0.00 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 76 s · en · evasnippets

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.246 before conversion and 0.908 after — it rose by 0.662. Neighbour-to-neighbour the worst pair went 0.246 → 0.936. (The earlier render, with segment 1 left raw, scores 0.806 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.204 in the original and +0.230 after conversion — 113 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Affection, -0.022 became -0.012.

Quality. Mean predicted overall quality across the segments went 2.96 → 3.11 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.246 → 0.908 +0.662identity cos neighbours 0.246 → 0.936d_b rescored +0.204 → +0.230d_a rescored -0.022 → -0.012d_a mined -0.022d_b mined 0.205min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_01_00_0000_6total 75.2schain gain +3.8 dBseam step 0.6 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, slightly bright, fairly smooth, balanced body, quiet background, energised, neutral tension, moderately variable
(affection, embarrassment, amusement · brisk, some disfluency, light breath, casual) And (ahem) um and I met Josh that night and he was this other girl's date and they were very much like, we're just friends, we're just friends, we're just friends. And I was like, okay. And that's why I have a whole joke now about my ex who left me for the girl he told me not to worry about because they kept saying they're just friends.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as affection, embarrassment, amusement; style: casual, playful; good recording, quiet background; genuineness 4.4/6; vocal-burst blend 6.1/10; 15.6s.
cond_podcastt_01_00_0000_681644_00125120 · in -20.8 dBFS · gain +0.8 dB · evasnippets-00316
(affection, embarrassment, amusement · brisk, some disfluency, light breath, casual) And (ahem) um and I met Josh that night and he was this other girl's date and they were very much like, we're just friends, we're just friends, we're just friends. And I was like, okay. And that's why I have a whole joke now about my ex who left me for the girl he told me not to worry about because they kept saying they're just friends.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as affection, embarrassment, amusement; style: casual, playful; good recording, quiet background; genuineness 4.4/6; vocal-burst blend 6.1/10; 15.6s.
cond_podcastt_01_00_0000_681644_00125120 · in -20.8 dBFS · gain +0.8 dB · evasnippets-00316
(interest, pleasure ecstasy, elation · normal-paced, frequent disfluency, normal breath, casual) Joe Coy. The huge, amazing, wonderful, talented Joe Coy. So Aurora, if she's played a really fun game where she's like, Travis, send me embarrassing pictures, and then I will send you gorgeous model shots.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as interest, pleasure ecstasy, elation; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.4/6; vocal-burst blend 3.2/10; 14.9s.
cond_podcastt_01_00_0000_681644_00137560 · in -18.9 dBFS · gain -1.1 dB · evasnippets-00316
(interest, pleasure ecstasy, elation · normal-paced, frequent disfluency, normal breath, casual) Joe Coy. The huge, amazing, wonderful, talented Joe Coy. So Aurora, if she's played a really fun game where she's like, Travis, send me embarrassing pictures, and then I will send you gorgeous model shots.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as interest, pleasure ecstasy, elation; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.4/6; vocal-burst blend 3.2/10; 14.9s.
cond_podcastt_01_00_0000_681644_00137560 · in -18.9 dBFS · gain -1.1 dB · evasnippets-00316
(interest, pleasure ecstasy, elation · normal-paced, frequent disfluency, normal breath, casual) Joe Coy. The huge, amazing, wonderful, talented Joe Coy. So Aurora, if she's played a really fun game where she's like, Travis, send me embarrassing pictures, and then I will send you gorgeous model shots.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as interest, pleasure ecstasy, elation; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.4/6; vocal-burst blend 3.2/10; 14.9s.
cond_podcastt_01_00_0000_681644_00137560 · in -18.9 dBFS · gain -1.1 dB · evasnippets-00316
Sexual Lust(unconstrained axis: Thankfulness Gratitude)identity +0.04 emotion 68 %   c-evasnippets-B1 · #7

This chain comes from the one-sided rule: only Sexual Lust had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Sexual Lust clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.25.

Nothing was asked of the other axis, and in fact Thankfulness Gratitude drifts down from 0.96 to 0.90 (-0.06), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.00 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 51 s · en · evasnippets

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.720 before conversion and 0.758 after — it rose by 0.038. Neighbour-to-neighbour the worst pair went 0.720 → 0.758. (The earlier render, with segment 1 left raw, scores 0.646 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sexual Lust moved +0.249 in the original and +0.169 after conversion — 68 % of the delta retained. On the other named axis, Thankfulness Gratitude, -0.060 became -0.062.

Quality. Mean predicted overall quality across the segments went 2.98 → 3.24 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.720 → 0.758 +0.038identity cos neighbours 0.720 → 0.758d_b rescored +0.249 → +0.169d_a rescored -0.060 → -0.062d_a mined -0.060d_b mined 0.249min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_00_02_0002_3total 50.0schain gain +4.1 dBseam step 1.5 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, slightly bright, good recording, no background noise, normal-paced, slightly relaxed, moderately variable, little disfluency
(thankfulness gratitude, contemplation, sadness · normally alert, average clarity, light breath, conversational) You know, you also mentioned something a little while ago about community. And here you were alone, and you only had so much help around you. Um, (low mumble) might that be number three, right?
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, little disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as thankfulness gratitude, contemplation, sadness; style: conversational, playful; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 0.8/10; 10.6s.
cond_podcastt_00_02_0002_332841_00138128 · in -24.0 dBFS · gain +4.0 dB · evasnippets-00321
(sexual lust, contentment, hope enthusiasm optimism · very low-energy, clear, normal breath, storytelling) whatever. And your body has wisdom. Your body will communicate with you if you can spend enough time, right, in quiet time, to hear, right? To hear the guidance that your body has for you. It's really amazing. I mean this is coming from the scientists and like I said, the neuroscientists.
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, slightly rough, full; clear, little disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as sexual lust, contentment, hope enthusiasm optimism; style: storytelling, playful; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 2.3/10; 19.9s.
cond_podcastt_00_02_0002_332841_00174736 · in -26.1 dBFS · gain +6.1 dB · evasnippets-00317
(sexual lust, contentment, hope enthusiasm optimism · very low-energy, clear, normal breath, storytelling) whatever. And your body has wisdom. Your body will communicate with you if you can spend enough time, right, in quiet time, to hear, right? To hear the guidance that your body has for you. It's really amazing. I mean this is coming from the scientists and like I said, the neuroscientists.
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, slightly rough, full; clear, little disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as sexual lust, contentment, hope enthusiasm optimism; style: storytelling, playful; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 2.3/10; 19.9s.
cond_podcastt_00_02_0002_332841_00174736 · in -26.1 dBFS · gain +6.1 dB · evasnippets-00317
Impatience and Irritability(unconstrained axis: Teasing)identity +0.01 emotion 111 %   c-evasnippets-B1 · #8

This chain comes from the one-sided rule: only Impatience and Irritability had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Impatience and Irritability clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Teasing drifts down from 0.99 to 0.90 (-0.09), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.00, then +0.00, then +0.24, then +0.00 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 70 s · en · evasnippets

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.882 before conversion and 0.896 after — it rose by 0.014. Neighbour-to-neighbour the worst pair went 0.882 → 0.917. (The earlier render, with segment 1 left raw, scores 0.783 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.236 in the original and +0.261 after conversion — 111 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Teasing, -0.091 became -0.042.

Quality. Mean predicted overall quality across the segments went 3.05 → 3.26 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.882 → 0.896 +0.014identity cos neighbours 0.882 → 0.917d_b rescored +0.236 → +0.261d_a rescored -0.091 → -0.042d_a mined -0.091d_b mined 0.235min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_00_01_0003_6total 69.2schain gain +3.2 dBseam step 0.3 dBcrossfades 100/100/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, slightly bright, fairly smooth, balanced body, average recording, quiet background, brisk, energised
(amusement, teasing, embarrassment · light breath, casual, conversational) Oh no, okay. Fair enough. Okay. No, on a personal level, Wes Craven. So, but other than Wes Craven, and I mean, and to be fair, the one we covered for Wes Craven wasn't exactly, you know, his top tier. But I did I do love Invitation to Hell. It was Invitation to Hell, right? That was.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as amusement, teasing, embarrassment; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.6/6; vocal-burst blend 6.4/10; 13.3s.
cond_podcastt_00_01_0003_639317_00023424 · in -18.0 dBFS · gain -2.0 dB · evasnippets-00321
(amusement, teasing, embarrassment · light breath, casual, conversational) Oh no, okay. Fair enough. Okay. No, on a personal level, Wes Craven. So, but other than Wes Craven, and I mean, and to be fair, the one we covered for Wes Craven wasn't exactly, you know, his top tier. But I did I do love Invitation to Hell. It was Invitation to Hell, right? That was.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as amusement, teasing, embarrassment; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.6/6; vocal-burst blend 6.4/10; 13.3s.
cond_podcastt_00_01_0003_639317_00023424 · in -18.0 dBFS · gain -2.0 dB · evasnippets-00321
(amusement, teasing, embarrassment · light breath, casual, conversational) Oh no, okay. Fair enough. Okay. No, on a personal level, Wes Craven. So, but other than Wes Craven, and I mean, and to be fair, the one we covered for Wes Craven wasn't exactly, you know, his top tier. But I did I do love Invitation to Hell. It was Invitation to Hell, right? That was.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as amusement, teasing, embarrassment; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.6/6; vocal-burst blend 6.4/10; 13.3s.
cond_podcastt_00_01_0003_639317_00023424 · in -18.0 dBFS · gain -2.0 dB · evasnippets-00321
(impatience and irritability, embarrassment, sourness · normal breath, casual, conversational) Yeah. But what but would you especially where the when I don't give anything away yet'cause I'm sure we'll spoil things by the end, but where it ends up and what she's willing to do, and you can say oh it's justified, totally. But what she does is raw. More I mean there's a more ethical point.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, slightly guarded; reads as impatience and irritability, embarrassment, sourness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.3/6; vocal-burst blend 5.8/10; 15.0s.
cond_podcastt_00_01_0003_639317_00078616 · in -18.0 dBFS · gain -2.0 dB · evasnippets-00316
(impatience and irritability, embarrassment, sourness · normal breath, casual, conversational) Yeah. But what but would you especially where the when I don't give anything away yet'cause I'm sure we'll spoil things by the end, but where it ends up and what she's willing to do, and you can say oh it's justified, totally. But what she does is raw. More I mean there's a more ethical point.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, slightly guarded; reads as impatience and irritability, embarrassment, sourness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.3/6; vocal-burst blend 5.8/10; 15.0s.
cond_podcastt_00_01_0003_639317_00078616 · in -18.0 dBFS · gain -2.0 dB · evasnippets-00316
Contentment(unconstrained axis: Awe)identity −0.05 emotion 86 %   c-evasnippets-B1 · #9

This chain comes from the one-sided rule: only Contentment had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Contentment clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.21.

Nothing was asked of the other axis, and in fact Awe drifts down from 0.96 to 0.47 (-0.48), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.21 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 33 s · en · evasnippets

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.666 before conversion and 0.618 after — it fell by 0.048. Neighbour-to-neighbour the worst pair went 0.666 → 0.618. (The earlier render, with segment 1 left raw, scores 0.453 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.209 in the original and +0.181 after conversion — 86 % of the delta retained, which is most of it. On the other named axis, Awe, -0.483 became +0.000.

Quality. Mean predicted overall quality across the segments went 3.10 → 3.17 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.666 → 0.618 -0.048identity cos neighbours 0.666 → 0.618d_b rescored +0.209 → +0.181d_a rescored -0.483 → +0.000d_a mined -0.483d_b mined 0.209min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_02_02_0003_9total 33.2schain gain +3.4 dBseam step 0.6 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, slightly bright, fairly smooth, balanced body, good recording, normal-paced, normally alert, slightly relaxed
(awe, interest, concentration · moderately variable, some disfluency, average clarity, casual) Okay, so we've got the rod, the reel. What about the line itself? Like is there anything special about fly fishing line? Oh yeah, the line is a whole system. You have the fly line, which is the weighted line that you cast. Then there's the leader, which is a tapered section that connects the fly line to your fly. It helps transfer energy smoothly and create a delicate presentation in the water. Wait, tapered? Why is that so important?
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as awe, interest, concentration; style: casual, conversational; good recording, quiet background; genuineness 1.6/6; vocal-burst blend 3.9/10; 23.7s.
cond_podcastt_02_02_0003_905766_00010376 · in -31.2 dBFS · gain +11.2 dB · evasnippets-00326
(contentment, sourness, teasing · fairly steady, little disfluency, clear, casual) Pretty much. You also have to think about your casting angle and how you're presenting your fly. You want it to land naturally and drift along without any drag, just like a real insect would.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contentment, sourness, teasing; style: casual, conversational; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.0/10; 9.6s.
cond_podcastt_02_02_0003_905766_00041744 · in -32.0 dBFS · gain +12.0 dB · evasnippets-00321
Anger(unconstrained axis: Bitterness)identity −0.03 emotion 83 %   c-evasnippets-B1 · #10

This chain comes from the one-sided rule: only Anger had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Anger strongly present — 0.78, higher than 78 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.20.

Nothing was asked of the other axis, and in fact Bitterness barely moves at all, sitting near 0.95 throughout.

It takes 5 clips to get there. Clip to clip the moves are +0.14, then -0.05, then +0.12, then -0.00 — not a clean run: step 2 moves back the other way by 0.05 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 124 s · en · evasnippets

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.946 before conversion and 0.915 after — it fell by 0.031. Neighbour-to-neighbour the worst pair went 0.940 → 0.927. (The earlier render, with segment 1 left raw, scores 0.705 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Anger moved +0.205 in the original and +0.171 after conversion — 83 % of the delta retained, which is most of it. On the other named axis, Bitterness, +0.031 became -0.008.

Quality. Mean predicted overall quality across the segments went 3.21 → 3.39 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.946 → 0.915 -0.031identity cos neighbours 0.940 → 0.927d_b rescored +0.205 → +0.171d_a rescored +0.031 → -0.008d_a mined 0.031d_b mined 0.204min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_02_02_0004_8total 122.4schain gain +2.8 dBseam step 0.3 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-bright, balanced body, quiet background, energised, some disfluency, average clarity
(bitterness, contemplation, triumph · normal-paced, neutral tension, moderately variable, dramatic) I know that again, y'all two have not ever, your parents have never had that. Perfection almost over there, right? (exhausted groan) The goal is justice. The attitude behind discipline is always love. Right? And its goal is the benefit and development of the person. So Hebrews 12 tells us about a displeased heavenly father who disciplines his children with a goal that's more corrective than punishment, but there's a punitive element in the discipline God hands out.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as bitterness, contemplation, triumph; style: dramatic, casual; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 2.4/10; 26.5s.
cond_podcastt_02_02_0004_893615_00129904 · in -24.3 dBFS · gain +4.3 dB · evasnippets-00154
(bitterness, jealousy and envy, malevolence malice · brisk, slightly relaxed, moderately variable, authoritative) It's a firm, loving parental chastisement to a child who's doing something wrong. Note the words that are used here. Scourging. When you think of a scourging, what does that necessarily mean? My daddy used to scourge me when I was a kid. I didn't realize that today. It's a whipping. Y'all ever had a whipping? Man, I tell you one thing. Daddy's whippings compared to God's whippings, I think I'd prefer my daddy's, right? Because they were done. And guys, they last a little bit. Chasting and chastisement. That's another word for correction. Bruh. Do y'all need anybody need some correction?
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, fairly guarded; reads as bitterness, jealousy and envy, malevolence malice; style: authoritative, dramatic; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 3.3/10; 28.4s.
cond_podcastt_02_02_0004_893615_00132552 · in -25.3 dBFS · gain +5.3 dB · evasnippets-00154
(disappointment, doubt, jealousy and envy · brisk, neutral tension, moderately variable, casual) I still need correction. I'm old. I still need it. I'm not there yet. Rebuke. I thought of rebuke, it means to check or restrain. You ever seen a dog on a collar? What's that doing? Checking him and restraining him. Do we like to be checked and restrained? I do not. I don't know why that is. I've never liked that. But all of these things that God does, He's doing this for our benefit.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, fairly guarded; reads as disappointment, doubt, jealousy and envy; style: casual, dramatic; average recording, quiet background; mildly explicit content; genuineness 4.3/6; vocal-burst blend 7.0/10; 23.4s.
cond_podcastt_02_02_0004_893615_00135392 · in -26.3 dBFS · gain +6.3 dB · evasnippets-00154
(malevolence malice, anger, awe · normal-paced, slightly relaxed, fairly steady, casual) And the Lord heard the sound of your words, and was angry, and took an oath, saying, Surely not one of these men of this evil generation shall see that good land of which I swore to give to your fathers, except Caleb the son of Jeff Jeffunah, he shall see it, and to him and his children I am giving the land on which he walked. Why? Because he wholly followed the Lord. The Lord was also angry with me for your sake, saving saying, Even you shall not go in there talking to Moses.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as malevolence malice, anger, awe; style: casual, monologue; good recording, quiet background; genuineness 2.0/6; vocal-burst blend 5.4/10; 24.3s.
cond_podcastt_02_02_0004_893615_00139376 · in -26.1 dBFS · gain +6.2 dB · evasnippets-00154
(anger, bitterness, shame · brisk, slightly relaxed, moderately variable, casual) God's still going to continue, no matter what you've done and turned away. So again, what does confession accomplish? Remember, confession, uh, (ahem) Winston and them aren't here, but confession is not about losing your salvation, right? We can't do that. We didn't obtain it, we can't do it. Are we renewing our justification? Both of those God does. Can't do it, right? Forgiveness and cleansing are two aspects of that promise.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as anger, bitterness, shame; style: casual, authoritative; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 4.1/10; 20.5s.
cond_podcastt_02_02_0004_893615_00150200 · in -25.7 dBFS · gain +5.7 dB · evasnippets-00154
Amusement(unconstrained axis: Infatuation)identity −0.02 emotion 98 %   c-evasnippets-B1 · #11

This chain comes from the one-sided rule: only Amusement had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Amusement strongly present — 0.78, higher than 78 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.20.

Nothing was asked of the other axis, and in fact Infatuation drifts down from 0.98 to 0.05 (-0.93), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.00, then +0.00, then +0.00 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 141 s · en · evasnippets

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.874 before conversion and 0.849 after — it fell by 0.025. Neighbour-to-neighbour the worst pair went 0.874 → 0.858. (The earlier render, with segment 1 left raw, scores 0.649 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Amusement moved +0.203 in the original and +0.199 after conversion — 98 % of the delta retained, which is essentially all of it. On the other named axis, Infatuation, -0.927 became -0.964.

Quality. Mean predicted overall quality across the segments went 3.11 → 3.34 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.874 → 0.849 -0.025identity cos neighbours 0.874 → 0.858d_b rescored +0.203 → +0.199d_a rescored -0.927 → -0.964d_a mined -0.927d_b mined 0.203min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_00_01_0000_8total 140.1schain gain +1.7 dBseam step 1.4 dBcrossfades 150/100/100/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(infatuation, contemplation, jealousy and envy · casual, conversational) So that's what I think. I think if he's with someone that's just kind of wanting to win a chip that's focused on that, like even like if he goes to another team, if he's with someone that's more like offensively minded, that wants to win the team like let's say if he goes to just this is just crazy, like best case scenario that the fucking (low mumble) um who just won who won the Super Bowl I can't, Mahomes.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as infatuation, contemplation, jealousy and envy; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.6/6; vocal-burst blend 10.0/10; 21.2s.
cond_podcastt_00_01_0000_879888_00168000 · in -29.4 dBFS · gain +9.4 dB · evasnippets-00317
(amusement, sexual lust, disappointment · casual, conversational) I don't really know. It was just one of those things where it was like, I get it. It was so w such an eye roll. And like it's just one of those things that was like, Come on, you know Yeah, we're here to buy it for Margot Robbie.'Cause like you didn't even try to cast anybody too I mean the biggest one for me this is just personal this is just personally for me was Huntress when they the actress they casted for Huntress. I felt like okay that drove me to see her. But then the other two, I was like, I wasn't I mean, they killed in the movie, but like I wasn't too like the star power wasn't there compared
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as amusement, sexual lust, disappointment; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.3/6; vocal-burst blend 10.0/10; 29.9s.
cond_podcastt_00_01_0000_879888_00197960 · in -25.8 dBFS · gain +5.8 dB · evasnippets-00298
(amusement, sexual lust, disappointment · casual, conversational) I don't really know. It was just one of those things where it was like, I get it. It was so w such an eye roll. And like it's just one of those things that was like, Come on, you know Yeah, we're here to buy it for Margot Robbie.'Cause like you didn't even try to cast anybody too I mean the biggest one for me this is just personal this is just personally for me was Huntress when they the actress they casted for Huntress. I felt like okay that drove me to see her. But then the other two, I was like, I wasn't I mean, they killed in the movie, but like I wasn't too like the star power wasn't there compared
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as amusement, sexual lust, disappointment; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.3/6; vocal-burst blend 10.0/10; 29.9s.
cond_podcastt_00_01_0000_879888_00197960 · in -25.8 dBFS · gain +5.8 dB · evasnippets-00298
(amusement, sexual lust, disappointment · casual, conversational) I don't really know. It was just one of those things where it was like, I get it. It was so w such an eye roll. And like it's just one of those things that was like, Come on, you know Yeah, we're here to buy it for Margot Robbie.'Cause like you didn't even try to cast anybody too I mean the biggest one for me this is just personal this is just personally for me was Huntress when they the actress they casted for Huntress. I felt like okay that drove me to see her. But then the other two, I was like, I wasn't I mean, they killed in the movie, but like I wasn't too like the star power wasn't there compared
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as amusement, sexual lust, disappointment; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.3/6; vocal-burst blend 10.0/10; 29.9s.
cond_podcastt_00_01_0000_879888_00197960 · in -25.8 dBFS · gain +5.8 dB · evasnippets-00298
(amusement, sexual lust, disappointment · casual, conversational) I don't really know. It was just one of those things where it was like, I get it. It was so w such an eye roll. And like it's just one of those things that was like, Come on, you know Yeah, we're here to buy it for Margot Robbie.'Cause like you didn't even try to cast anybody too I mean the biggest one for me this is just personal this is just personally for me was Huntress when they the actress they casted for Huntress. I felt like okay that drove me to see her. But then the other two, I was like, I wasn't I mean, they killed in the movie, but like I wasn't too like the star power wasn't there compared
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as amusement, sexual lust, disappointment; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.3/6; vocal-burst blend 10.0/10; 29.9s.
cond_podcastt_00_01_0000_879888_00197960 · in -25.8 dBFS · gain +5.8 dB · evasnippets-00298
Interest(unconstrained axis: Infatuation)identity +0.58 emotion 92 %   c-evasnippets-B1 · #12

This chain comes from the one-sided rule: only Interest had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Interest clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.39.

Nothing was asked of the other axis, and in fact Infatuation barely moves at all, sitting near 0.99 throughout.

It takes 4 clips to get there. Clip to clip the moves are +0.15, then +0.23, then +0.00 — a plateau around step 3, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 83 s · en · evasnippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.301 before conversion and 0.876 after — it rose by 0.575. Neighbour-to-neighbour the worst pair went 0.308 → 0.888. (The earlier render, with segment 1 left raw, scores 0.706 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.385 in the original and +0.356 after conversion — 92 % of the delta retained, which is essentially all of it. On the other named axis, Infatuation, -0.006 became -0.012.

Quality. Mean predicted overall quality across the segments went 3.10 → 3.22 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.301 → 0.876 +0.575identity cos neighbours 0.308 → 0.888d_b rescored +0.385 → +0.356d_a rescored -0.006 → -0.012d_a mined -0.006d_b mined 0.385min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_02_00_0000_1total 82.5schain gain +4.3 dBseam step 0.6 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, slightly bright, fairly smooth, good recording, neutral tension, moderately variable, average clarity, wide pitch range
(infatuation, shame, embarrassment · normal-paced, very low-energy, frequent disfluency, casual) (low mumble) Um, this was just when I was graduating college, I was really starting to realize I might be a little gay. I was letting those thoughts come into my mind. I had tried dating boys in college, and I found myself in this community of people who loved Fifth Harmony and were also gay. And
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as infatuation, shame, embarrassment; style: casual, monologue; good recording, no background noise; genuineness 3.4/6; vocal-burst blend 5.0/10; 22.3s.
cond_podcastt_02_00_0000_154122_00130559 · in -25.2 dBFS · gain +5.2 dB · evasnippets-00303
(contentment, elation, shame · brisk, normally alert, some disfluency, casual) Like a cyber cafe. Yeah, like I did make friends. Like and this was my outlet. I could be straight grace walking through the world, but when I was in my phone on Twitter talking to these friends, I had I was out. So it was about Fifth Harmony, let me tell you. Like I knew all their songs. I voted for every award they were up for.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as contentment, elation, shame; style: casual, conversational; good recording, quiet background; genuineness 4.9/6; vocal-burst blend 9.4/10; 19.7s.
cond_podcastt_02_00_0000_154122_00134400 · in -22.5 dBFS · gain +2.5 dB · evasnippets-00280
(interest, infatuation, amusement · brisk, energised, some disfluency, casual) I think that like I when I do look at those videos, yeah, they're slowed down and romanticized, but like what is that Carly Kloss, Taylor Swift photo of them? They look like they're kissing on the mouth, like holding each other's faces. Like that you look like you're making out. You look like Karita and I in in high school in those pictures. Yeah. Oh yeah. And I mean there was this whole thing where Carly Kloss came to one of the LA shows.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as interest, infatuation, amusement; style: casual, conversational; good recording, no background noise; genuineness 3.5/6; vocal-burst blend 6.1/10; 20.5s.
cond_podcastt_02_00_0000_154122_00182200 · in -24.2 dBFS · gain +4.2 dB · evasnippets-00280
(interest, infatuation, amusement · brisk, energised, some disfluency, casual) I think that like I when I do look at those videos, yeah, they're slowed down and romanticized, but like what is that Carly Kloss, Taylor Swift photo of them? They look like they're kissing on the mouth, like holding each other's faces. Like that you look like you're making out. You look like Karita and I in in high school in those pictures. Yeah. Oh yeah. And I mean there was this whole thing where Carly Kloss came to one of the LA shows.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as interest, infatuation, amusement; style: casual, conversational; good recording, no background noise; genuineness 3.5/6; vocal-burst blend 6.1/10; 20.5s.
cond_podcastt_02_00_0000_154122_00182200 · in -24.2 dBFS · gain +4.2 dB · evasnippets-00280
Relief(unconstrained axis: Disappointment)identity +0.08 emotion 135 %   c-evasnippets-B1 · #13

This chain comes from the one-sided rule: only Relief had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Relief strongly present — 0.78, higher than 78 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.21.

Nothing was asked of the other axis, and in fact Disappointment drifts down from 1.00 to 0.39 (-0.60), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.13, then +0.07, then +0.01, then +0.00 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 137 s · es · evasnippets

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.757 before conversion and 0.835 after — it rose by 0.077. Neighbour-to-neighbour the worst pair went 0.738 → 0.824. (The earlier render, with segment 1 left raw, scores 0.797 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.214 in the original and +0.290 after conversion — 135 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Disappointment, -0.604 became -0.595.

Quality. Mean predicted overall quality across the segments went 3.07 → 3.30 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.757 → 0.835 +0.077identity cos neighbours 0.738 → 0.824d_b rescored +0.214 → +0.290d_a rescored -0.604 → -0.595d_a mined -0.604d_b mined 0.214min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang esspeaker cond_podcastt_00_03_0004_6total 135.7schain gain +2.7 dBseam step 1.6 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-bright, quiet background, light breath
(disappointment, bitterness, jealousy and envy · fast, energised, neutral tension, monologue) Reportage que pone en tela de juicio (low mumble) la probidad o la honestidad de algunos candidatos. ¿Crees tú que veamos esto también in las campañas judiciales? Porque en las campañas políticas tradicionales it's more common. Pero en las judiciales, pues estamos viendo, ¿no? Yo creo que hay mucho que hacer en este mes que nos queda, ¿no? (low mumble)
full caption & clip details
A young adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as disappointment, bitterness, jealousy and envy; style: monologue, dramatic; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 8.8/10; 24.2s.
cond_podcastt_00_03_0004_678600_00080668 · in -38.2 dBFS · gain +18.2 dB · evasnippets-00012
(interest, bitterness, sourness · fast, energised, neutral tension, monologue) Probably a good idea, but the execution has been disposable and the execution has been a little, not a contrary, and also the organization, the information, etc. How to be a reform (low mumble) that we audit (low mumble) this proposal of reform judicial and of election of the judgment in particular.
full caption & clip details
A child masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as interest, bitterness, sourness; style: monologue, storytelling; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 9.2/10; 25.6s.
cond_podcastt_00_03_0004_678600_00127224 · in -37.4 dBFS · gain +17.4 dB · evasnippets-00012
(disappointment, relief, jealousy and envy · measured, very low-energy, neutral tension, monologue) Morris, ya para ir cerrando, tres lecciones positivas que nos deja la elección judicial o que nos están dejando las campañas judiciales. Tres lecciones positivas. Qué difficile pregunta. No, yo creo que una lección positiva es que la gente se dé cuenta que nuestro trabajo is muy importante (ahem) como asesores incommunication. Yo creo que la gente tiene el derecho.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as disappointment, relief, jealousy and envy; style: monologue, casual; below-average recording, quiet background; genuineness 3.1/6; vocal-burst blend 4.5/10; 29.8s.
cond_podcastt_00_03_0004_678600_00165416 · in -38.9 dBFS · gain +18.9 dB · evasnippets-00289
(relief, thankfulness gratitude, affection · fast, subdued, slightly relaxed, monologue) Voy a regresar un poco. (ahem) Voy a pedirte a ver, a manera de ejercicio, calificación para cada uno de los actores involucrados, digamos, en este proceso. Calificación del poder ejecutivo. ¿Cuánto le pondrías tú al poder ejecutivo en cuanto a su intervención o no dentro de este proceso electoral? Yo creo que más que el poder ejecutivo, la presidenta, ha sido (low mumble) muy activista.
full caption & clip details
A young adult masculine voice; delivery is subdued, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as relief, thankfulness gratitude, affection; style: monologue, casual; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 8.8/10; 28.4s.
cond_podcastt_00_03_0004_678600_00188504 · in -38.3 dBFS · gain +18.3 dB · evasnippets-00327
(relief, thankfulness gratitude, affection · fast, subdued, slightly relaxed, monologue) Voy a regresar un poco. (ahem) Voy a pedirte a ver, a manera de ejercicio, calificación para cada uno de los actores involucrados, digamos, en este proceso. Calificación del poder ejecutivo. ¿Cuánto le pondrías tú al poder ejecutivo en cuanto a su intervención o no dentro de este proceso electoral? Yo creo que más que el poder ejecutivo, la presidenta, ha sido (low mumble) muy activista.
full caption & clip details
A young adult masculine voice; delivery is subdued, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as relief, thankfulness gratitude, affection; style: monologue, casual; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 8.8/10; 28.4s.
cond_podcastt_00_03_0004_678600_00188504 · in -38.3 dBFS · gain +18.3 dB · evasnippets-00327
Anger(unconstrained axis: Hope Enthusiasm Optimism)identity −0.01 emotion 96 %   c-evasnippets-B1 · #14

This chain comes from the one-sided rule: only Anger had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Anger clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Hope Enthusiasm Optimism barely moves at all, sitting near 1.00 throughout.

It takes 3 clips to get there. Clip to clip the moves are +0.02, then +0.22 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 58 s · en · evasnippets

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.862 before conversion and 0.850 after — it fell by 0.012. Neighbour-to-neighbour the worst pair went 0.873 → 0.837. (The earlier render, with segment 1 left raw, scores 0.644 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Anger moved +0.242 in the original and +0.233 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Hope Enthusiasm Optimism, -0.044 became -0.018.

Quality. Mean predicted overall quality across the segments went 2.93 → 3.12 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.862 → 0.850 -0.012identity cos neighbours 0.873 → 0.837d_b rescored +0.242 → +0.233d_a rescored -0.044 → -0.018d_a mined -0.044d_b mined 0.242min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_02_02_0002_3total 57.8schain gain +3.1 dBseam step 1.2 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · slightly bright, fairly smooth, average recording, quiet background, brisk, energised, variable, some disfluency
(hope enthusiasm optimism, doubt, fear · neutral tension, somewhat unclear, casual, conversational) I need his guidance. I need him to show me. Lord, I need you. Oh, I need you. That there's a song that goes like that, and I just pray it. Lord, I need you. Oh, I need you. And so I'm like kind of standing on this scripture that says I I need to build myself up in my most holy faith by praying in the Holy Spirit. And I'll start praying. I'll pray naturally. I'll pray in my heavenly language. I'll just pray, pray, pray.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism, doubt, fear; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.6/6; vocal-burst blend 9.0/10; 23.4s.
cond_podcastt_02_02_0002_328751_00208440 · in -25.8 dBFS · gain +5.8 dB · evasnippets-00161
(pleasure ecstasy, fatigue exhaustion, hope enthusiasm optimism · slightly tense, average clarity, casual, storytelling) And it's amazing how that span of time gets a little bit longer, a little bit longer, and a little bit longer every day. And then I'm running out of time to do stuff like in the morning because I'm spending so much time praying for praying to God in the morning. And I it's beautiful. It's beautiful.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly tense, variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as pleasure ecstasy, fatigue exhaustion, hope enthusiasm optimism; style: casual, storytelling; average recording, quiet background; mildly explicit content; genuineness 4.0/6; vocal-burst blend 5.6/10; 15.8s.
cond_podcastt_02_02_0002_328751_00210776 · in -25.4 dBFS · gain +5.4 dB · evasnippets-00280
(anger, impatience and irritability, hope enthusiasm optimism · slightly tense, very clear, ranting, dramatic) But it doesn't fire. You know that? You know that's what happens? It's not as flammable, it's not as combustive, it doesn't have the same properties that are explosive. I thought that that's a pretty good explanation, right there. You put we put natural, we put logic, we put our own thinking into our lives.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly tense, variable; timbre is slightly cool, slightly bright, fairly smooth, thin; very clear, some disfluency, wide pitch range, normal breath; affect is negative, very dominant, fairly guarded; reads as anger, impatience and irritability, hope enthusiasm optimism; style: ranting, dramatic; average recording, quiet background; mildly explicit content; genuineness 3.1/6; vocal-burst blend 2.8/10; 19.0s.
cond_podcastt_02_02_0002_328751_00290704 · in -20.8 dBFS · gain +0.8 dB · evasnippets-00280
Hope Enthusiasm Optimism(unconstrained axis: Confusion)identity +0.05 emotion 93 %   c-evasnippets-B1 · #15

This chain comes from the one-sided rule: only Hope Enthusiasm Optimism had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Hope Enthusiasm Optimism strongly present — 0.77, higher than 77 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.23.

Nothing was asked of the other axis, and in fact Confusion drifts down from 0.98 to 0.49 (-0.50), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.00, then +0.00, then +0.00, then +0.23 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 76 s · en · evasnippets

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.672 before conversion and 0.718 after — it rose by 0.045. Neighbour-to-neighbour the worst pair went 0.672 → 0.739. (The earlier render, with segment 1 left raw, scores 0.620 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.233 in the original and +0.217 after conversion — 93 % of the delta retained, which is essentially all of it. On the other named axis, Confusion, -0.495 became -0.351.

Quality. Mean predicted overall quality across the segments went 2.99 → 3.15 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.672 → 0.718 +0.045identity cos neighbours 0.672 → 0.739d_b rescored +0.233 → +0.217d_a rescored -0.495 → -0.351d_a mined -0.495d_b mined 0.232min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_01_00_0000_5total 74.7schain gain +4.1 dBseam step 2.6 dBcrossfades 100/100/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, slightly bright, fairly smooth, balanced body, moderately variable, wide pitch range, light breath
(confusion, amusement, pleasure ecstasy · normal-paced, normally alert, neutral tension, casual) not answer and block him on up everything. I've never blocked anyone. But when you say like that, I would probably be like, Hey, like I can't today. And then if he says like, oh what about tomorrow? I think I'd probably leave him on Yeah, I'd probably leave him on R. You give them one answer and if they don't get the hit and they text Bob
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as confusion, amusement, pleasure ecstasy; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 4.5/6; vocal-burst blend 5.1/10; 16.9s.
cond_podcastt_01_00_0000_508163_00205527 · in -23.8 dBFS · gain +3.8 dB · evasnippets-00327
(confusion, amusement, pleasure ecstasy · normal-paced, normally alert, neutral tension, casual) not answer and block him on up everything. I've never blocked anyone. But when you say like that, I would probably be like, Hey, like I can't today. And then if he says like, oh what about tomorrow? I think I'd probably leave him on Yeah, I'd probably leave him on R. You give them one answer and if they don't get the hit and they text Bob
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as confusion, amusement, pleasure ecstasy; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 4.5/6; vocal-burst blend 5.1/10; 16.9s.
cond_podcastt_01_00_0000_508163_00205527 · in -23.8 dBFS · gain +3.8 dB · evasnippets-00327
(confusion, amusement, pleasure ecstasy · normal-paced, normally alert, neutral tension, casual) not answer and block him on up everything. I've never blocked anyone. But when you say like that, I would probably be like, Hey, like I can't today. And then if he says like, oh what about tomorrow? I think I'd probably leave him on Yeah, I'd probably leave him on R. You give them one answer and if they don't get the hit and they text Bob
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as confusion, amusement, pleasure ecstasy; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 4.5/6; vocal-burst blend 5.1/10; 16.9s.
cond_podcastt_01_00_0000_508163_00205527 · in -23.8 dBFS · gain +3.8 dB · evasnippets-00327
(confusion, amusement, pleasure ecstasy · normal-paced, normally alert, neutral tension, casual) not answer and block him on up everything. I've never blocked anyone. But when you say like that, I would probably be like, Hey, like I can't today. And then if he says like, oh what about tomorrow? I think I'd probably leave him on Yeah, I'd probably leave him on R. You give them one answer and if they don't get the hit and they text Bob
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as confusion, amusement, pleasure ecstasy; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 4.5/6; vocal-burst blend 5.1/10; 16.9s.
cond_podcastt_01_00_0000_508163_00205527 · in -23.8 dBFS · gain +3.8 dB · evasnippets-00327
(hope enthusiasm optimism, elation, thankfulness gratitude · brisk, energised, slightly tense, conversational) Okay, guys, so thank you so much for tuning into episode four of Cheers to That. We really hope you enjoyed this week's topic.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism, elation, thankfulness gratitude; style: conversational, casual; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 1.7/10; 7.6s.
cond_podcastt_01_00_0000_508163_00242232 · in -21.2 dBFS · gain +1.2 dB · evasnippets-00320
Teasing(unconstrained axis: Shame)identity +0.14 emotion 22 %   c-evasnippets-B1 · #16

This chain comes from the one-sided rule: only Teasing had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Teasing strongly present — 0.78, higher than 78 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.21.

Nothing was asked of the other axis, and in fact Shame drifts down from 0.98 to 0.25 (-0.73), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.21, then +0.00, then +0.00 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 76 s · en · evasnippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.731 before conversion and 0.873 after — it rose by 0.143. Neighbour-to-neighbour the worst pair went 0.731 → 0.881. (The earlier render, with segment 1 left raw, scores 0.609 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Teasing moved +0.211 in the original and +0.046 after conversion — 22 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Shame, -0.733 became -0.190.

Quality. Mean predicted overall quality across the segments went 3.04 → 3.34 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.731 → 0.873 +0.143identity cos neighbours 0.731 → 0.881d_b rescored +0.211 → +0.046d_a rescored -0.733 → -0.190d_a mined -0.733d_b mined 0.211min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_00_02_0002_9total 74.7schain gain +4.5 dBseam step 0.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, slightly rough, balanced body, quiet background, normal-paced, average clarity, light breath
(shame, fear, confusion · very low-energy, slightly relaxed, fairly steady, casual) Finch, near Automata, Hellblade Senua Sacrifice, Wolfenstein 2, and Horizon Zero Dawn. Outstanding story telling and narrative development in a game. Okay, what remains of Edith Finch? Never heard of it. Hellblade Senua's Sacrifice, kind of heard of it, but again, not really look much into it. Wolfenstein 2, beyond some of the marketing I've seen.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as shame, fear, confusion; style: casual, monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 4.7/10; 26.4s.
cond_podcastt_00_02_0002_969803_00150408 · in -27.1 dBFS · gain +7.1 dB · evasnippets-00327
(teasing, hope enthusiasm optimism, amusement · energised, neutral tension, moderately variable, casual) Fair enough. Yeah, no, they're all skilled. I will say this though award for the best eye color for a voice actor has to go to Brian Bloom. He's got them gorgeous hazel eyes. I tell you what. I tell you what. Alright, Robbie, take us down our next category.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as teasing, hope enthusiasm optimism, amusement; style: casual, playful; good recording, quiet background; genuineness 4.6/6; vocal-burst blend 3.4/10; 16.3s.
cond_podcastt_00_02_0002_969803_00293768 · in -24.1 dBFS · gain +4.1 dB · evasnippets-00316
(teasing, hope enthusiasm optimism, amusement · energised, neutral tension, moderately variable, casual) Fair enough. Yeah, no, they're all skilled. I will say this though award for the best eye color for a voice actor has to go to Brian Bloom. He's got them gorgeous hazel eyes. I tell you what. I tell you what. Alright, Robbie, take us down our next category.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as teasing, hope enthusiasm optimism, amusement; style: casual, playful; good recording, quiet background; genuineness 4.6/6; vocal-burst blend 3.4/10; 16.3s.
cond_podcastt_00_02_0002_969803_00293768 · in -24.1 dBFS · gain +4.1 dB · evasnippets-00316
(teasing, hope enthusiasm optimism, amusement · energised, neutral tension, moderately variable, casual) Fair enough. Yeah, no, they're all skilled. I will say this though award for the best eye color for a voice actor has to go to Brian Bloom. He's got them gorgeous hazel eyes. I tell you what. I tell you what. Alright, Robbie, take us down our next category.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as teasing, hope enthusiasm optimism, amusement; style: casual, playful; good recording, quiet background; genuineness 4.6/6; vocal-burst blend 3.4/10; 16.3s.
cond_podcastt_00_02_0002_969803_00293768 · in -24.1 dBFS · gain +4.1 dB · evasnippets-00316
Interest(unconstrained axis: Intoxication Altered States of Consciousness)identity +0.30 emotion 102 %   c-evasnippets-B1 · #17

This chain comes from the one-sided rule: only Interest had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Interest strongly present — 0.77, higher than 77 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.22.

Nothing was asked of the other axis, and in fact Intoxication Altered States of Consciousness drifts down from 0.98 to 0.03 (-0.95), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.16, then +0.06 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 81 s · en · evasnippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.453 before conversion and 0.751 after — it rose by 0.298. Neighbour-to-neighbour the worst pair went 0.370 → 0.816. (The earlier render, with segment 1 left raw, scores 0.628 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.224 in the original and +0.229 after conversion — 102 % of the delta retained, which is essentially all of it. On the other named axis, Intoxication Altered States of Consciousness, -0.954 became -0.961.

Quality. Mean predicted overall quality across the segments went 3.11 → 3.34 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.453 → 0.751 +0.298identity cos neighbours 0.370 → 0.816d_b rescored +0.224 → +0.229d_a rescored -0.954 → -0.961d_a mined -0.954d_b mined 0.224min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_00_02_0005_4total 79.8schain gain +3.4 dBseam step 1.2 dBcrossfades 100/150/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, fairly steady, moderate pitch range
(intoxication altered states of consciousness, doubt, fatigue exhaustion · measured, subdued, relaxed, casual) I agree, Jim. I actually think it's (low mumble) uh I think if it comes out right now Mahomes is playing. I don't think the line moves more than definitely not more than two and a half to the Chiefs. And I think if it's said right now that uh (low mumble) Mahomes is out, I think that line goes all the way up to three or four to the Bengals.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as intoxication altered states of consciousness, doubt, fatigue exhaustion; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 5.6/10; 16.3s.
cond_podcastt_00_02_0005_463998_00206680 · in -30.3 dBFS · gain +10.3 dB · evasnippets-00316
(intoxication altered states of consciousness, doubt, fatigue exhaustion · measured, subdued, relaxed, casual) I agree, Jim. I actually think it's (low mumble) uh I think if it comes out right now Mahomes is playing. I don't think the line moves more than definitely not more than two and a half to the Chiefs. And I think if it's said right now that uh (low mumble) Mahomes is out, I think that line goes all the way up to three or four to the Bengals.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as intoxication altered states of consciousness, doubt, fatigue exhaustion; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 5.6/10; 16.3s.
cond_podcastt_00_02_0005_463998_00206680 · in -30.3 dBFS · gain +10.3 dB · evasnippets-00316
(hope enthusiasm optimism, amusement, jealousy and envy · normal-paced, normally alert, neutral tension, casual) quickly. But again, I think Mahomes is gonna be a little bit more than Which says two things how great of a quarterback Patrick Mahomes is, and how great of a team this Bengals team might actually be, and how we've written them off as the third fiddle to the third wheel to the Bills and the Chiefs. And last year they made the Super Bowl and it was like, oh yeah, okay, like let's see it again and now they're doing it again. And it's like this team might actually be the best team in the AFC the last two years and we're just writing it off.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism, amusement, jealousy and envy; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 9.6/10; 28.8s.
cond_podcastt_00_02_0005_463998_00209008 · in -29.7 dBFS · gain +9.7 dB · evasnippets-00034
(interest, triumph, contentment · normal-paced, normally alert, neutral tension, casual) I mean, I think you could say they've they've been the best team in the NFL for the past two years anyway, right? I mean, the Rams fell off the face of the earth this year, who would have they were the only team that were better than them last year 'cause they beat him in the Super Bowl. And here's Cincinnati again, back to back AFC championships. And um (low mumble) Joe Burrow is the second best quarterback in the NFL. It's him and Mahom.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, triumph, contentment; style: casual, conversational; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 9.0/10; 18.9s.
cond_podcastt_00_02_0005_463998_00214112 · in -28.4 dBFS · gain +8.4 dB · evasnippets-00325
Contemplation(unconstrained axis: Amusement)identity −0.00 emotion 145 %   c-evasnippets-B1 · #18

This chain comes from the one-sided rule: only Contemplation had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Contemplation strongly present — 0.78, higher than 78 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.21.

Nothing was asked of the other axis, and in fact Amusement drifts down from 0.98 to 0.82 (-0.16), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.16, then +0.06, then +0.00 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 98 s · en · evasnippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.913 before conversion and 0.909 after — it fell by 0.004. Neighbour-to-neighbour the worst pair went 0.883 → 0.894. (The earlier render, with segment 1 left raw, scores 0.676 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.215 in the original and +0.312 after conversion — 145 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Amusement, -0.159 became -0.063.

Quality. Mean predicted overall quality across the segments went 3.06 → 3.28 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.913 → 0.909 -0.004identity cos neighbours 0.883 → 0.894d_b rescored +0.215 → +0.312d_a rescored -0.159 → -0.063d_a mined -0.159d_b mined 0.215min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_00_02_0004_3total 97.4schain gain +4.2 dBseam step 0.3 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(amusement, elation, interest · moderately variable, average clarity, casual, conversational) And the reason why I wanted to stop there was because of Jim Morrison's grave. Like the I even like we were like, where are we gonna stop in between? They're like, we can stop in Paris. And I was like, what the hell's in Paris besides, you know, like why would I want to go to Paris with with a bunch of guys, you know, kind of thing? So my friend was like, Jim Morrison's grave is in Paris. And I was like, all right, that's amazing. Like, let's go. So we went, we were in Paris for a night. But like, that's the thing. You start doing these things that, you know, people
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as amusement, elation, interest; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 10.0/10; 24.2s.
cond_podcastt_00_02_0004_36363_00116232 · in -42.9 dBFS · gain +22.9 dB · evasnippets-00325
(fatigue exhaustion, helplessness, relief · fairly steady, somewhat unclear, casual, conversational) Sometimes you just have to, you know, pack the bag and go and do it. I I've said that from day one, you know, you don't know how long you're gonna last. So try to get through it as much as you can, save up the money and go and do it. And even if you don't get the full effect, like I was in London, like one of my biggest things I wanted to go to London. Like I was in London for literally a day by myself and I just had to walk around because you know I didn't want to get lost or anything.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as fatigue exhaustion, helplessness, relief; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 10.0/10; 21.9s.
cond_podcastt_00_02_0004_36363_00119240 · in -44.7 dBFS · gain +24.7 dB · evasnippets-00325
(contemplation, longing, shame · moderately variable, somewhat unclear, casual, conversational) ex yeah, we didn't have guns in the house. And even my dad being part of the, you know, National Guard, he was a firefighter. So, you know, I don't think I've ever had I don't think I remember ever having a gun or or anything like that in the house. (low mumble) Um and most of the time, so if you think about it, like I I'm not really hanging out with my sisters. So I just spent a lot of time alone. Like most of my, you know, kind of thinking back to what I did as a kid, I got myself into not trouble.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as contemplation, longing, shame; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 10.0/10; 25.9s.
cond_podcastt_00_02_0004_36363_00184384 · in -41.9 dBFS · gain +21.9 dB · evasnippets-00029
(contemplation, longing, shame · moderately variable, somewhat unclear, casual, conversational) ex yeah, we didn't have guns in the house. And even my dad being part of the, you know, National Guard, he was a firefighter. So, you know, I don't think I've ever had I don't think I remember ever having a gun or or anything like that in the house. (low mumble) Um and most of the time, so if you think about it, like I I'm not really hanging out with my sisters. So I just spent a lot of time alone. Like most of my, you know, kind of thinking back to what I did as a kid, I got myself into not trouble.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as contemplation, longing, shame; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 10.0/10; 25.9s.
cond_podcastt_00_02_0004_36363_00184384 · in -41.9 dBFS · gain +21.9 dB · evasnippets-00029
Sourness(unconstrained axis: Teasing)identity +0.01 emotion 332 %   c-evasnippets-B1 · #19

This chain comes from the one-sided rule: only Sourness had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Sourness strongly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Teasing drifts down from 0.97 to 0.90 (-0.07), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.00 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 42 s · en · evasnippets

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.791 before conversion and 0.797 after — it rose by 0.005. Neighbour-to-neighbour the worst pair went 0.791 → 0.819. (The earlier render, with segment 1 left raw, scores 0.656 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sourness moved +0.238 in the original and +0.790 after conversion — 332 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Teasing, -0.069 became -0.009.

Quality. Mean predicted overall quality across the segments went 3.25 → 3.32 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.791 → 0.797 +0.005identity cos neighbours 0.791 → 0.819d_b rescored +0.238 → +0.790d_a rescored -0.069 → -0.009d_a mined -0.070d_b mined 0.237min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_00_00_0005_7total 41.9schain gain +1.5 dBseam step 0.7 dBcrossfades 100/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, normal-paced, normally alert, some disfluency, average clarity
(teasing, anger, impatience and irritability · neutral tension, moderately variable, wide pitch range, casual) And rather than just be like, Is police brutality a problem? Yes. Should we do something about police brutality? Yeah. You know, like rather than Yeah now it's like F you and F you and now
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as teasing, anger, impatience and irritability; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.2/6; vocal-burst blend 9.2/10; 11.3s.
cond_podcastt_00_00_0005_784463_00307416 · in -28.6 dBFS · gain +8.6 dB · evasnippets-00324
(sourness, anger, thankfulness gratitude · slightly relaxed, fairly steady, moderate pitch range, casual) Like he he just we can have a reasonable disagreement on what the actual best solution for sick people in the United States is. Yeah. And that and that's where like I'd say coming full circle, like the comedy club should be the one room in America. It's not a church, it's not a classroom, it's not a
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as sourness, anger, thankfulness gratitude; style: casual, conversational; average recording, no background noise; mildly explicit content; genuineness 4.3/6; vocal-burst blend 8.7/10; 15.4s.
cond_podcastt_00_00_0005_784463_00324544 · in -27.4 dBFS · gain +7.4 dB · evasnippets-00327
(sourness, anger, thankfulness gratitude · slightly relaxed, fairly steady, moderate pitch range, casual) Like he he just we can have a reasonable disagreement on what the actual best solution for sick people in the United States is. Yeah. And that and that's where like I'd say coming full circle, like the comedy club should be the one room in America. It's not a church, it's not a classroom, it's not a
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as sourness, anger, thankfulness gratitude; style: casual, conversational; average recording, no background noise; mildly explicit content; genuineness 4.3/6; vocal-burst blend 8.7/10; 15.4s.
cond_podcastt_00_00_0005_784463_00324544 · in -27.4 dBFS · gain +7.4 dB · evasnippets-00327
Interest(unconstrained axis: Bitterness)identity +0.22 emotion 61 %   c-evasnippets-B1 · #20

This chain comes from the one-sided rule: only Interest had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Interest strongly present — 0.78, higher than 78 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.20.

Nothing was asked of the other axis, and in fact Bitterness drifts down from 0.98 to 0.85 (-0.13), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.00 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the evasnippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 75 s · en · evasnippets

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.527 before conversion and 0.751 after — it rose by 0.224. Neighbour-to-neighbour the worst pair went 0.527 → 0.808. (The earlier render, with segment 1 left raw, scores 0.406 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.201 in the original and +0.121 after conversion — 61 % of the delta retained. On the other named axis, Bitterness, -0.133 became -0.105.

Quality. Mean predicted overall quality across the segments went 2.93 → 3.31 (+0.38) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.527 → 0.751 +0.224identity cos neighbours 0.527 → 0.808d_b rescored +0.201 → +0.121d_a rescored -0.133 → -0.105d_a mined -0.133d_b mined 0.201min_cos_consec (site) —min_cos_anchor (site) —dataset evasnippetslang enspeaker cond_podcastt_00_00_0003_8total 74.0schain gain +3.3 dBseam step 0.6 dBcrossfades 100/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, some disfluency
(bitterness, contemplation, anger · neutral tension, moderately variable, wide pitch range, casual) Like still 60, 70 years out, we still can't know the truth, you know, all those people involved are gone. Like uh (low mumble) uh what's the still like that truth has to be scarier than just certain people. That that's what it tells me. That truth has to be bigger than just certain people were doing fuckery. Yeah.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as bitterness, contemplation, anger; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.6/6; vocal-burst blend 8.6/10; 18.8s.
cond_podcastt_00_00_0003_898115_00134936 · in -20.6 dBFS · gain +0.6 dB · evasnippets-00323
(interest, astonishment surprise, hope enthusiasm optimism · slightly relaxed, fairly steady, moderate pitch range, casual) Before you were even born. Like that was what's startling about watching Oliver Stone's film and then seeing it take shape in the Congress, like uh (low mumble) uh (low mumble) injuriate lawmakers and they change the law, uh, (ahem) but still like over 30 years. All right, we'll give you the records, but we're gonna do it slowly over 30 years. And they do that, and people have been camped out in their hall of records. Oliver Stone made a follow-up documentary called JFK Revisited, that is the best JFK documentary I've ever seen because it
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, astonishment surprise, hope enthusiasm optimism; style: casual, monologue; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 8.8/10; 27.7s.
cond_podcastt_00_00_0003_898115_00138368 · in -20.5 dBFS · gain +0.5 dB · evasnippets-00324
(interest, astonishment surprise, hope enthusiasm optimism · slightly relaxed, fairly steady, moderate pitch range, casual) Before you were even born. Like that was what's startling about watching Oliver Stone's film and then seeing it take shape in the Congress, like uh (low mumble) uh (low mumble) injuriate lawmakers and they change the law, uh, (ahem) but still like over 30 years. All right, we'll give you the records, but we're gonna do it slowly over 30 years. And they do that, and people have been camped out in their hall of records. Oliver Stone made a follow-up documentary called JFK Revisited, that is the best JFK documentary I've ever seen because it
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, astonishment surprise, hope enthusiasm optimism; style: casual, monologue; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 8.8/10; 27.7s.
cond_podcastt_00_00_0003_898115_00138368 · in -20.5 dBFS · gain +0.5 dB · evasnippets-00324