This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_c-snippets-PXR.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words
are a real non-speech sound, printed where it happens.
The full generated caption for any clip is under “full caption & clip
details”. Its perceived-gender and background-noise clauses were re-rendered from
the numeric buckets, because the versions stored in the corpus index had those two
ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
The tags and captions describe the ORIGINAL clips — they are the
corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the
converted audio are the ones printed in each card's own paragraph and chips.
This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Emotional Numbness clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.27.
At the same time Infatuation goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.42. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.12, then +0.15 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.07 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.00 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.07, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 29 s · snippets
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.058 before conversion and 0.670 after — it rose by 0.612. Neighbour-to-neighbour the worst pair went -0.007 → 0.594. (The earlier render, with segment 1 left raw, scores 0.558 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.263 in the original and +0.124 after conversion — 47 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Infatuation, -0.453 became -0.343.
Quality. Mean predicted overall quality across the segments went 2.91 → 3.07 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.058 → 0.670+0.612identity cos neighbours -0.007 → 0.594d_b rescored +0.263 → +0.124d_a rescored -0.453 → -0.343d_a mined -0.424d_b mined 0.272min_cos_consec (site) -0.0038min_cos_anchor (site) 0.0706dataset snippetslang ?speaker batch214_part4_batch214_patotal 28.2schain gain +2.9 dBseam step 0.5 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult somewhat feminine voice · good recording, normally alert, slightly relaxed
(infatuation, affection, longing · normal-paced, fairly steady, some disfluency, casual)And I said, (low mumble) uh, my name and who I was related to and who he was related to, because he the family was a first cousin.
full caption & clip details
A young adult somewhat feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as infatuation, affection, longing; style: casual, conversational; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 2.4/10; 9.3s.
batch214_part4_batch214_part4_chunk_3_1_73922 · in -14.1 dBFS · gain -5.9 dB · snippets-00594
(slow, steady, frequent disfluency, monologue)globally that generates about a sixth of Ford's global automotive revenue. This strike includes
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is warm, slightly dark, slightly rough, full; very clear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, fairly guarded; no dominant emotion; style: monologue, storytelling; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.0/10; 9.2s.
batch214_part4_batch214_part4_chunk_3_1_74107 · in -20.0 dBFS · gain -0.0 dB · snippets-00594
(normal-paced, steady, no disfluency, narration)Drawing from the teachings of Ninjitsu, Chloe embodies the principle of hiding in plain sight, an unsuspecting flower company executive by day, and a relentless vigilante by night.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.1s.
batch214_part4_batch214_part4_chunk_3_1_74240 · in -22.4 dBFS · gain +2.5 dB · snippets-00594
This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Infatuation clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.40.
At the same time Hope Enthusiasm Optimism goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.63 (higher than 63 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.14, then +0.24, then +0.02 — a plateau around step 3, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.35 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.13 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.35, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 30 s · snippets
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.433 before conversion and 0.743 after — it rose by 0.310. Neighbour-to-neighbour the worst pair went 0.113 → 0.667. (The earlier render, with segment 1 left raw, scores 0.608 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.371 in the original and +0.669 after conversion — 180 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Hope Enthusiasm Optimism, -0.294 became -0.294.
Quality. Mean predicted overall quality across the segments went 2.72 → 2.92 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.433 → 0.743+0.310identity cos neighbours 0.113 → 0.667d_b rescored +0.371 → +0.669d_a rescored -0.294 → -0.294d_a mined -0.300d_b mined 0.397min_cos_consec (site) 0.1271min_cos_anchor (site) 0.3512dataset snippetslang ?speaker batch2_part3_batch2_part3_total 29.0schain gain +2.8 dBseam step 0.4 dBcrossfades 150/150/100 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, fairly steady, average clarity
(hope enthusiasm optimism · fully relaxed, some disfluency, casual, conversational)and he's going to be one of those guys. I think Nate's legacy in freestyle motocross is bringing a professionalism.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as hope enthusiasm optimism; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.0/6; vocal-burst blend 6.2/10; 6.4s.
batch2_part3_batch2_part3_chunk_1010_1_1017301 · in -33.0 dBFS · gain +13.0 dB · snippets-01045
(sadness, contentment, contemplation·neutral tension, frequent disfluency, casual, whispered)And I know that's what keeps him really centered and really narrowed with what what the choices are that he makes. You know, I think Nate and I hit it off so well because we have the same faith, we're both Christian and you know, we I think God plays a huge, Jesus plays a huge emphasis in our lives.
full caption & clip details
An adult somewhat feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as sadness, contentment, contemplation; style: casual, whispered; average recording, quiet background; mildly explicit content; genuineness 4.0/6; vocal-burst blend 3.3/10; 15.5s.
batch2_part3_batch2_part3_chunk_1010_1_1017336 · in -33.4 dBFS · gain +13.4 dB · snippets-01045
(infatuation, contentment, embarrassment·slightly relaxed, frequent disfluency, casual, conversational)really (ahem) is. I'd describe Nate as (low mumble) um, just a really nice, humble guy.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as infatuation, contentment, embarrassment; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 3.3/10; 4.4s.
batch2_part3_batch2_part3_chunk_1010_1_1017386 · in -33.1 dBFS · gain +13.1 dB · snippets-01045
(infatuation, pride, triumph· slightly relaxed, little disfluency, casual, conversational)That was my second contest I've ever ridden professionally.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as infatuation, pride, triumph; style: casual, conversational; good recording, no background noise; genuineness 3.3/6; vocal-burst blend 3.2/10; 3.2s.
batch2_part3_batch2_part3_chunk_1010_1_1017418 · in -32.6 dBFS · gain +12.7 dB · snippets-01045
This chain comes from the proxy rule: the same two-sided test as above, but because Fatigue Exhaustion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Fatigue Exhaustion around average — 0.57, higher than 57 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.30.
At the same time Contemplation goes the other way, from 0.82 (higher than 82 % of clips in this corpus) to 0.55 (higher than 55 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.10, then +0.20 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.75 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.75 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.75, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 23 s · snippets
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.733 before conversion and 0.627 after — it fell by 0.106. Neighbour-to-neighbour the worst pair went 0.733 → 0.712. (The earlier render, with segment 1 left raw, scores 0.383 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.296 in the original and +0.401 after conversion — 136 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contemplation, -0.272 became -0.168.
Quality. Mean predicted overall quality across the segments went 2.66 → 2.98 (+0.32) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.733 → 0.627-0.106identity cos neighbours 0.733 → 0.712d_b rescored +0.296 → +0.401d_a rescored -0.272 → -0.168d_a mined -0.271d_b mined 0.296min_cos_consec (site) 0.7546min_cos_anchor (site) 0.7546dataset snippetslang ?speaker batch23_part0_batch23_parttotal 22.4schain gain +4.0 dBseam step 0.4 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, quiet background, normally alert, slightly relaxed, fairly steady, moderate pitch range
(measured, some disfluency, somewhat unclear, casual)graphic novel ideas that we've transitioned over into (ahem) visual novels. So, uh, (low mumble)
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 1.8/10; 6.1s.
batch23_part0_batch23_part0_chunk_1205_1_1123248 · in -22.3 dBFS · gain +2.3 dB · snippets-00727
(distress, helplessness, fear·normal-paced, some disfluency, average clarity, casual)You're going to have failures a lot. If you're the type of person that collapses and has a has a breakdown every time there's a failure,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as distress, helplessness, fear; style: casual, monologue; good recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.5/10; 8.4s.
batch23_part0_batch23_part0_chunk_1205_1_1123295 · in -22.5 dBFS · gain +2.5 dB · snippets-00727
(measured, frequent disfluency, average clarity, casual)getting close to half. And (ahem) it morphs the music and uh (low mumble) film industry, (ahem) uh the gaming industry more morphs.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 3.5/6; vocal-burst blend 0.8/10; 8.2s.
batch23_part0_batch23_part0_chunk_1205_1_1123312 · in -22.4 dBFS · gain +2.4 dB · snippets-00727
This chain comes from the proxy rule: the same two-sided test as above, but because Fatigue Exhaustion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Fatigue Exhaustion around average — 0.42, lower than 58 % of clips in this corpus — and ends with it clearly present at 0.71, higher than 71 % of clips in this corpus. That is a total rise of 0.29.
At the same time Malevolence Malice goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.05, then +0.24 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.91 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 23 s · snippets
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.878 before conversion and 0.844 after — it fell by 0.034. Neighbour-to-neighbour the worst pair went 0.878 → 0.859. (The earlier render, with segment 1 left raw, scores 0.738 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.265 in the original and +0.171 after conversion — 65 % of the delta retained. On the other named axis, Malevolence Malice, -0.261 became -0.333.
Quality. Mean predicted overall quality across the segments went 2.70 → 2.92 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.878 → 0.844-0.034identity cos neighbours 0.878 → 0.859d_b rescored +0.265 → +0.171d_a rescored -0.261 → -0.333d_a mined -0.261d_b mined 0.290min_cos_consec (site) 0.9092min_cos_anchor (site) 0.8997dataset snippetslang ?speaker batch278_part4_batch278_patotal 22.3schain gain +3.1 dBseam step 0.1 dBcrossfades 100/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normally alert, slightly relaxed, fairly steady
(brisk, almost no disfluency, clear, formal)The herbicide Orange was destroyed by incineration aboard a German incinerator ship, MV Vulcanus, in 1977.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, didactic; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.8/10; 6.7s.
batch278_part4_batch278_part4_chunk_975_1_1189051 · in -18.6 dBFS · gain -1.4 dB · snippets-00931
(normal-paced, almost no disfluency, clear, narration)Marine Jim Divine described the changes to Sand Island in 1944. Sand Island at the time consisted of two islands connected by a single lane coral roadway about 600 feet long.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.3/10; 9.2s.
batch278_part4_batch278_part4_chunk_975_1_1189074 · in -21.5 dBFS · gain +1.5 dB · snippets-00931
(brisk, little disfluency, average clarity, playful)One of the island's anti-aircraft batteries have been destined for Wake Island before the island fell to the Japanese on December 23rd, 1941.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: playful, authoritative; good recording, quiet background; genuineness 1.9/6; vocal-burst blend 1.7/10; 6.6s.
batch278_part4_batch278_part4_chunk_975_1_1189242 · in -19.8 dBFS · gain -0.2 dB · snippets-00931
This chain comes from the proxy rule: the same two-sided test as above, but because Thankfulness Gratitude is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Thankfulness Gratitude clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.31.
At the same time Contempt goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.09, then +0.22 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.15 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.15 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.15, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 18 s · snippets
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.149 before conversion and 0.288 after — it rose by 0.139. Neighbour-to-neighbour the worst pair went 0.149 → 0.350. (The earlier render, with segment 1 left raw, scores 0.192 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.314 in the original and +0.255 after conversion — 81 % of the delta retained, which is most of it. On the other named axis, Contempt, -0.375 became -0.371.
Quality. Mean predicted overall quality across the segments went 2.85 → 2.89 (+0.03) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.149 → 0.288+0.139identity cos neighbours 0.149 → 0.350d_b rescored +0.314 → +0.255d_a rescored -0.375 → -0.371d_a mined -0.376d_b mined 0.314min_cos_consec (site) 0.1483min_cos_anchor (site) 0.1483dataset snippetslang ?speaker batch135_part1_batch135_patotal 17.0schain gain +1.8 dBseam step 2.0 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, clear, moderate pitch range
(contempt, disgust, bitterness · measured, steady, almost no disfluency, formal)Like everyone else, indigenous children need certain books, not any books. Here's why.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, disgust, bitterness; style: formal, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 1.0/10; 6.0s.
batch135_part1_batch135_part1_chunk_2212_1_2084823 · in -18.4 dBFS · gain -1.6 dB · snippets-00185
(confusion, doubt· measured, fairly steady, some disfluency, didactic)But he also has some questions for Ju Galang. What if Cao Cao does not come my way?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as confusion, doubt; style: didactic, formal; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 0.8/10; 6.0s.
batch135_part1_batch135_part1_chunk_2212_1_2084843 · in -22.1 dBFS · gain +2.0 dB · snippets-00185
(thankfulness gratitude·brisk, fairly steady, no disfluency, casual)This episode was produced by Brett Bachman and Rebecca Ramirez, and edited by Andrea Kissick.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as thankfulness gratitude; style: casual, authoritative; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.0/10; 5.3s.
batch135_part1_batch135_part1_chunk_2212_1_2084919 · in -21.4 dBFS · gain +1.4 dB · snippets-00185
This chain comes from the proxy rule: the same two-sided test as above, but because Fatigue Exhaustion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Fatigue Exhaustion clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.31.
At the same time Doubt goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.35. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.12 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.13 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.08 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.13, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 10 s · snippets
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.128 before conversion and 0.443 after — it rose by 0.315. Neighbour-to-neighbour the worst pair went 0.064 → 0.416. (The earlier render, with segment 1 left raw, scores 0.237 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.309 in the original and +0.154 after conversion — 50 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Doubt, -0.354 became -0.079.
Quality. Mean predicted overall quality across the segments went 2.21 → 2.68 (+0.48) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.128 → 0.443+0.315identity cos neighbours 0.064 → 0.416d_b rescored +0.309 → +0.154d_a rescored -0.354 → -0.079d_a mined -0.354d_b mined 0.310min_cos_consec (site) 0.0832min_cos_anchor (site) 0.1317dataset snippetslang ?speaker batch73_part3_batch73_parttotal 9.6schain gain +0.5 dBseam step 5.1 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, quiet background, normally alert, fully relaxed, fairly steady, moderate pitch range
(doubt · measured, frequent disfluency, slurred, casual)(ahem) um and (low mumble) uh he talked about, you know,
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, fully relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; slurred, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: casual, conversational; poor recording, quiet background; genuineness 4.6/6; vocal-burst blend 1.2/10; 3.1s.
batch73_part3_batch73_part3_chunk_1658_1_1546962 · in -27.2 dBFS · gain +7.2 dB · snippets-01267
(disgust, pain, helplessness·normal-paced, frequent disfluency, slurred, casual)And like killing themselves or anything. Girls? No.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, fairly steady; timbre is neutral-toned, dark, slightly rough, slightly thin; slurred, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as disgust, pain, helplessness; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 3.7/10; 3.2s.
batch73_part3_batch73_part3_chunk_1658_1_1547030 · in -28.0 dBFS · gain +8.0 dB · snippets-01267
(fatigue exhaustion, emotional numbness· normal-paced, some disfluency, somewhat unclear, casual)yera ma kipe (ahem) (low mumble) agying wana nabatigitsi
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as fatigue exhaustion, emotional numbness; style: casual, conversational; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 2.6/10; 3.6s.
batch73_part3_batch73_part3_chunk_1658_1_1547165 · in -18.7 dBFS · gain -1.3 dB · snippets-01267
This chain comes from the proxy rule: the same two-sided test as above, but because Sourness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Sourness below average — 0.28, lower than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.62.
At the same time Infatuation goes the other way, from 0.92 (higher than 92 % of clips in this corpus) to 0.42 (lower than 58 % of clips in this corpus), a change of -0.50. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.16, then +0.23 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.06 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.27 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.06, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 15 s · snippets
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.062 before conversion and 0.516 after — it rose by 0.454. Neighbour-to-neighbour the worst pair went 0.269 → 0.474. (The earlier render, with segment 1 left raw, scores 0.428 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sourness moved +0.583 in the original and +0.639 after conversion — 109 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Infatuation, -0.501 became -0.347.
Quality. Mean predicted overall quality across the segments went 2.54 → 2.68 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.062 → 0.516+0.454identity cos neighbours 0.269 → 0.474d_b rescored +0.583 → +0.639d_a rescored -0.501 → -0.347d_a mined -0.501d_b mined 0.617min_cos_consec (site) 0.2739min_cos_anchor (site) 0.0647dataset snippetslang ?speaker batch27_part3_batch27_parttotal 14.1schain gain +1.7 dBseam step 1.9 dBcrossfades 150/150/100 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady
(infatuation · normal-paced, some disfluency, average clarity, conversational)And (low mumble) uh, but beyond that, you never heard about White Canvas trip again.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation; style: conversational, casual; good recording, no background noise; genuineness 2.8/6; vocal-burst blend 2.6/10; 3.7s.
batch27_part3_batch27_part3_chunk_1236_1_885437 · in -18.2 dBFS · gain -1.8 dB · snippets-00940
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; style: casual, conversational; average recording, no background noise; no dominant emotion; genuineness 4.2/6; vocal-burst blend 1.7/10; 3.2s.
batch27_part3_batch27_part3_chunk_1236_1_885592 · in -23.0 dBFS · gain +3.0 dB · snippets-00940
(normal-paced, some disfluency, average clarity, casual)posted on Twitter saying, does anyone have any information on this person's name?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual; good recording, no background noise; genuineness 3.5/6; vocal-burst blend 2.5/10; 3.8s.
batch27_part3_batch27_part3_chunk_1236_1_885829 · in -27.4 dBFS · gain +7.4 dB · snippets-00940
(normal-paced, some disfluency, average clarity, casual)George Soros doesn't even hold a candlestick to Peter Thiel in terms of (low mumble) uh.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; explicit content; genuineness 3.4/6; vocal-burst blend 2.1/10; 3.9s.
batch27_part3_batch27_part3_chunk_1236_1_885953 · in -24.6 dBFS · gain +4.5 dB · snippets-00940
Intoxication Altered States of Consciousness ↓ / Concentration ↑identity +0.11emotion 106 % c-snippets-PXR · #8
This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Concentration clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.31.
At the same time Intoxication Altered States of Consciousness goes the other way, from 0.75 (higher than 75 % of clips in this corpus) to 0.49 (right about the corpus median), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are -0.01, then +0.12, then +0.20 — not a clean run: step 1 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.70 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.77 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.70, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 31 s · snippets
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.677 before conversion and 0.784 after — it rose by 0.107. Neighbour-to-neighbour the worst pair went 0.737 → 0.828. (The earlier render, with segment 1 left raw, scores 0.571 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.310 in the original and +0.329 after conversion — 106 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Intoxication Altered States of Consciousness, -0.256 became -0.271.
Quality. Mean predicted overall quality across the segments went 2.99 → 3.04 (+0.05) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.677 → 0.784+0.107identity cos neighbours 0.737 → 0.828d_b rescored +0.310 → +0.329d_a rescored -0.256 → -0.271d_a mined -0.256d_b mined 0.312min_cos_consec (site) 0.7748min_cos_anchor (site) 0.6978dataset snippetslang ?speaker batch79_part0_batch79_parttotal 29.7schain gain +1.2 dBseam step 1.0 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, light breath
(steady, almost no disfluency, clear, monologue)That is the root cause and because of that he was skeptical to accept Sita.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 1.5/10; 5.0s.
batch79_part0_batch79_part0_chunk_1704_1_1535438 · in -19.1 dBFS · gain -0.9 dB · snippets-01293
(pain·fairly steady, some disfluency, average clarity, monologue)And that is what people have always been doing with regards to Uttara Ramayan. And the reason for this is their strong emotional connect.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pain; style: monologue, casual; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 2.2/10; 7.5s.
batch79_part0_batch79_part0_chunk_1704_1_1535454 · in -16.9 dBFS · gain -3.1 dB · snippets-01293
(steady, little disfluency, clear, monologue)So that is Skanda Puranam for you, standing as a strong argument that Uttara Kanda or Uttara Ramayanam is authentic and is for real.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.9/10; 8.5s.
batch79_part0_batch79_part0_chunk_1704_1_1535701 · in -17.3 dBFS · gain -2.7 dB · snippets-01293
(concentration, emotional numbness, contemplation·fairly steady, some disfluency, somewhat unclear, monologue)That is the reason people often jump to the conclusions that Uttara Ramayanam is very contentious or controversial and that is the reason it is a prakship, it is not part of Ramayan.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, emotional numbness, contemplation; style: monologue; average recording, no background noise; genuineness 3.3/6; vocal-burst blend 1.4/10; 9.2s.
batch79_part0_batch79_part0_chunk_1704_1_1535805 · in -21.2 dBFS · gain +1.2 dB · snippets-01293
This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Emotional Numbness around average — 0.47, lower than 53 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.40.
At the same time Concentration goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.58 (higher than 58 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.23 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.07 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.07 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.07, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 22 s · snippets
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.077 before conversion and 0.596 after — it rose by 0.519. Neighbour-to-neighbour the worst pair went 0.077 → 0.444. (The earlier render, with segment 1 left raw, scores 0.526 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.399 in the original and +0.123 after conversion — 31 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Concentration, -0.382 became -0.400.
Quality. Mean predicted overall quality across the segments went 2.71 → 2.94 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.077 → 0.596+0.519identity cos neighbours 0.077 → 0.444d_b rescored +0.399 → +0.123d_a rescored -0.382 → -0.400d_a mined -0.382d_b mined 0.399min_cos_consec (site) 0.0673min_cos_anchor (site) 0.0673dataset snippetslang ?speaker batch76_part0_batch76_parttotal 21.9schain gain +3.7 dBseam step 2.1 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, normally alert, slightly relaxed, moderate pitch range, light breath
(concentration · measured, steady, almost no disfluency, narration)If you take the measurement of the base perimeter of the great pyramid, go along and measure each of its four sides and add that together.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: narration, formal; good recording, no background noise; explicit content; genuineness 1.3/6; vocal-burst blend 0.9/10; 7.4s.
batch76_part0_batch76_part0_chunk_167_1_363035 · in -35.6 dBFS · gain +15.6 dB · snippets-01278
(awe·brisk, fairly steady, almost no disfluency, newsreading)lithic building with gigantic sculpted stones is considered by almost all the experts to have begun around 2,500 BC. From the Great Pyramid of Giza to Stonehenge in England.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as awe; style: newsreading, narration; good recording, no background noise; mildly explicit content; genuineness 0.4/6; vocal-burst blend 1.0/10; 11.3s.
batch76_part0_batch76_part0_chunk_167_1_363113 · in -32.3 dBFS · gain +12.3 dB · snippets-01278
(slow, fairly steady, some disfluency, casual)And so they existed on the peripheries of France, of
full caption & clip details
A young adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, whispered; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 3.0/10; 3.5s.
batch76_part0_batch76_part0_chunk_167_1_363140 · in -30.6 dBFS · gain +10.6 dB · snippets-01278
This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Emotional Numbness around average — 0.56, higher than 56 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.34.
At the same time Concentration goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.12, then +0.22 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores -0.14 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.05 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (-0.14, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 27 s · snippets
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.137 before conversion and 0.418 after — it rose by 0.555. Neighbour-to-neighbour the worst pair went -0.048 → 0.513. (The earlier render, with segment 1 left raw, scores 0.192 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
The emotional move did not survive. Re-scored end to end, Emotional Numbness moved +0.334 in the original and -0.136 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Concentration, -0.265 became -0.315.
Quality. Mean predicted overall quality across the segments went 3.05 → 3.16 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 -0.137 → 0.418+0.555identity cos neighbours -0.048 → 0.513d_b rescored +0.334 → -0.136d_a rescored -0.265 → -0.315d_a mined -0.264d_b mined 0.338min_cos_consec (site) -0.0497min_cos_anchor (site) -0.1447dataset snippetslang ?speaker batch108_part3_batch108_patotal 26.4schain gain +2.0 dBseam step 0.7 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · fairly smooth
(concentration · measured, normally alert, slightly relaxed, didactic)For this I'm going to use this data which has X1 to X 19 variable and the Y variable.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: didactic, monologue; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.8/10; 8.6s.
batch108_part3_batch108_part3_chunk_1970_1_2189316 · in -23.6 dBFS · gain +3.5 dB · snippets-00041
(relief, pride·normal-paced, normally alert, neutral tension, conversational)for the world. And so you can probably tell that (ahem) that myself and many other aid workers kind of view the net benefit as pretty limited in relation to what's going on.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as relief, pride; style: conversational, casual; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 6.3/10; 8.9s.
batch108_part3_batch108_part3_chunk_1970_1_2189343 · in -28.9 dBFS · gain +8.9 dB · snippets-00041
(slow, very low-energy, slightly relaxed, monologue)Some, more than, increased by, plus, and together, all mean addition.
full caption & clip details
A child feminine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is cool, slightly bright, fairly smooth, very thin; crisply articulate, almost no disfluency, very wide pitch range, minimal breath; affect is positive, neutral stance, neutral openness; no dominant emotion; style: monologue, didactic; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.8/10; 9.3s.
batch108_part3_batch108_part3_chunk_1970_1_2189482 · in -25.6 dBFS · gain +5.6 dB · snippets-00041
This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Contemplation around average — 0.57, higher than 57 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.33.
At the same time Affection goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.18 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.25 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.25 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.25, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 24 s · snippets
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.212 before conversion and 0.772 after — it rose by 0.560. Neighbour-to-neighbour the worst pair went 0.212 → 0.755. (The earlier render, with segment 1 left raw, scores 0.606 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.332 in the original and +0.292 after conversion — 88 % of the delta retained, which is most of it. On the other named axis, Affection, -0.314 became -0.883.
Quality. Mean predicted overall quality across the segments went 2.70 → 2.82 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.212 → 0.772+0.560identity cos neighbours 0.212 → 0.755d_b rescored +0.332 → +0.292d_a rescored -0.314 → -0.883d_a mined -0.314d_b mined 0.330min_cos_consec (site) 0.2495min_cos_anchor (site) 0.2495dataset snippetslang ?speaker batch48_part2_batch48_parttotal 23.4schain gain +2.5 dBseam step 1.2 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · fairly smooth, normal-paced, normally alert
(affection, interest · slightly relaxed, fairly steady, some disfluency, casual)a fleet of (ahem) um thinking, a lead if you will of thinking women. She is participatory in politics. She was president of the Republicans Colored Women's League. She's a part of the national.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as affection, interest; style: casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 2.5/10; 13.4s.
batch48_part2_batch48_part2_chunk_1428_1_1598057 · in -21.9 dBFS · gain +1.9 dB · snippets-01135
(slightly relaxed, fairly steady, some disfluency, casual)Can you talk a little bit about that and talk about the importance?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 2.9/6; vocal-burst blend 2.2/10; 3.8s.
batch48_part2_batch48_part2_chunk_1428_1_1598270 · in -18.0 dBFS · gain -2.0 dB · snippets-01135
(neutral tension, moderately variable, frequent disfluency, casual)of how she reads the period of enslavement. So what she's doing is changing
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: casual, storytelling; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 1.6/10; 6.5s.
batch48_part2_batch48_part2_chunk_1428_1_1598323 · in -18.9 dBFS · gain -1.1 dB · snippets-01135
Sexual Lust ↓ / Intoxication Altered States of Consciousness ↑identity +0.02emotion 111 % c-snippets-PXR · #12
This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Intoxication Altered States of Consciousness around average — 0.53, higher than 53 % of clips in this corpus — and ends with it strongly present at 0.79, higher than 79 % of clips in this corpus. That is a total rise of 0.26.
At the same time Sexual Lust goes the other way, from 0.70 (higher than 70 % of clips in this corpus) to 0.44 (lower than 56 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.04, then +0.22 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.71 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.74 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.71, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 13 s · snippets
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.736 before conversion and 0.758 after — it rose by 0.022. Neighbour-to-neighbour the worst pair went 0.745 → 0.741. (The earlier render, with segment 1 left raw, scores 0.727 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.239 in the original and +0.266 after conversion — 111 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Sexual Lust, -0.258 became -0.295.
Quality. Mean predicted overall quality across the segments went 2.64 → 2.74 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.736 → 0.758+0.022identity cos neighbours 0.745 → 0.741d_b rescored +0.239 → +0.266d_a rescored -0.258 → -0.295d_a mined -0.258d_b mined 0.257min_cos_consec (site) 0.7360min_cos_anchor (site) 0.7137dataset snippetslang ?speaker batch272_part3_batch272_patotal 12.6schain gain +2.1 dBseam step 1.0 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(formal, narration)To do so, first professional divers were hired to blast bombs underwater.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.9/10; 4.9s.
batch272_part3_batch272_part3_chunk_918_1_1036433 · in -23.4 dBFS · gain +3.4 dB · snippets-00901
(narration, monologue)For the south side, the hard strata was 50 feet below the seabed level.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: narration, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 2.4/10; 4.4s.
batch272_part3_batch272_part3_chunk_918_1_1036690 · in -24.2 dBFS · gain +4.2 dB · snippets-00901
(formal, monologue)Now, the length of the unsupported bridge deck is reduced.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.9/10; 3.6s.
batch272_part3_batch272_part3_chunk_918_1_1036753 · in -23.2 dBFS · gain +3.2 dB · snippets-00901
This chain comes from the proxy rule: the same two-sided test as above, but because Pleasure Ecstasy is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Pleasure Ecstasy clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.36.
At the same time Sexual Lust goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.10, then +0.18, then +0.08 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.41 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.41 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.41, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 53 s · snippets
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.446 before conversion and 0.645 after — it rose by 0.199. Neighbour-to-neighbour the worst pair went 0.474 → 0.635. (The earlier render, with segment 1 left raw, scores 0.593 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pleasure Ecstasy moved +0.360 in the original and +0.697 after conversion — 193 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Sexual Lust, -0.336 became -0.234.
Quality. Mean predicted overall quality across the segments went 2.86 → 3.08 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.446 → 0.645+0.199identity cos neighbours 0.474 → 0.635d_b rescored +0.360 → +0.697d_a rescored -0.336 → -0.234d_a mined -0.341d_b mined 0.360min_cos_consec (site) 0.4056min_cos_anchor (site) 0.4056dataset snippetslang ?speaker batch33_part0_batch33_parttotal 52.4schain gain +3.1 dBseam step 3.1 dBcrossfades 100/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, quiet background, average clarity, light breath
(sexual lust, sourness, amusement · brisk, energised, neutral tension, casual)Like and it she there needs to be some like the thing underneath has to be it it was the same red vinyl latex so you couldn't tell that something changed.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sexual lust, sourness, amusement; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 8.0/10; 8.4s.
batch33_part0_batch33_part0_chunk_1297_1_1090726 · in -21.8 dBFS · gain +1.8 dB · snippets-01062
(infatuation, sexual lust, amusement ·normal-paced, normally alert, slightly relaxed, casual)uh, (low mumble) that she got to sing when you got to mama on on television.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as infatuation, sexual lust, amusement; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 3.8/10; 4.8s.
batch33_part0_batch33_part0_chunk_1297_1_1090911 · in -24.2 dBFS · gain +4.2 dB · snippets-01062
(confusion, astonishment surprise, amusement ·brisk, energised, neutral tension, casual)I was, I don't know why Kevin Bacon was the one that really took me out. I was like, this is the last person I expected to see on, I know why he's there. I was just
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as confusion, astonishment surprise, amusement; style: casual, conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 9.2/10; 8.2s.
batch33_part0_batch33_part0_chunk_1297_1_1090935 · in -22.6 dBFS · gain +2.6 dB · snippets-01062
(pleasure ecstasy, pride, hope enthusiasm optimism· brisk, energised, slightly relaxed, casual)Now Alma has a diverse network of therapists to fit your unique needs. In the easy-to-use directory, you can filter for gender, sexual orientation, race, etcetera. Every therapist at Alma has a detailed profile so you can get a better picture of their working style, expertise, and even why they pursued a career in mental health. And it's sometimes important to make sure that your, in fact it is always important to make sure that your therapist is compatible with you and learning about their experience helps a lot. I think it is so important to take care of your mental health and I love how easy Alma makes it.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as pleasure ecstasy, pride, hope enthusiasm optimism; style: casual, conversational; good recording, quiet background; genuineness 2.2/6; vocal-burst blend 7.1/10; 31.5s.
batch33_part0_batch33_part0_chunk_1297_1_1091115 · in -28.5 dBFS · gain +8.5 dB · snippets-01062
This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Contemplation clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.30.
At the same time Concentration goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.05 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.70 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.72 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.70, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 21 s · snippets
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.606 before conversion and 0.770 after — it rose by 0.164. Neighbour-to-neighbour the worst pair went 0.677 → 0.771. (The earlier render, with segment 1 left raw, scores 0.571 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.299 in the original and +0.290 after conversion — 97 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.376 became -0.429.
Quality. Mean predicted overall quality across the segments went 2.89 → 3.03 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.606 → 0.770+0.164identity cos neighbours 0.677 → 0.771d_b rescored +0.299 → +0.290d_a rescored -0.376 → -0.429d_a mined -0.377d_b mined 0.299min_cos_consec (site) 0.7188min_cos_anchor (site) 0.7013dataset snippetslang ?speaker batch71_part3_batch71_parttotal 20.3schain gain +2.5 dBseam step 0.8 dBcrossfades 100/100 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normal-paced, normally alert, average clarity
(concentration · slightly relaxed, fairly steady, some disfluency, authoritative)So we randomized people between two conditions. They either get the (ahem) meditation while they're undergoing ultraviolet light or they just get the ultraviolet light by itself.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: authoritative, conversational; good recording, quiet background; genuineness 3.7/6; vocal-burst blend 2.5/10; 8.9s.
batch71_part3_batch71_part3_chunk_1636_1_1902236 · in -19.4 dBFS · gain -0.6 dB · snippets-01257
(contemplation·neutral tension, moderately variable, frequent disfluency, conversational)(ahem) Uh, if you stare at that word for too long, it doesn't mean anything, as you know. But, (ahem) uh, I want to make a distinc- (ahem) a distinction between
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contemplation; style: conversational, casual; good recording, quiet background; genuineness 4.3/6; vocal-burst blend 4.3/10; 8.0s.
batch71_part3_batch71_part3_chunk_1636_1_1902334 · in -21.3 dBFS · gain +1.3 dB · snippets-01257
(contemplation, doubt·slightly relaxed, fairly steady, frequent disfluency, casual)so when the doing comes out of being I think you know then (surprised gasp)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as contemplation, doubt; style: casual, conversational; good recording, no background noise; genuineness 4.1/6; vocal-burst blend 4.3/10; 3.6s.
batch71_part3_batch71_part3_chunk_1636_1_1902368 · in -19.7 dBFS · gain -0.3 dB · snippets-01257
Intoxication Altered States of Consciousness ↓ / Infatuation ↑identity +0.46emotion 28 % c-snippets-PXR · #15
This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Infatuation below average — 0.33, lower than 67 % of clips in this corpus — and ends with it strongly present at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.49.
At the same time Intoxication Altered States of Consciousness goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.47 (lower than 53 % of clips in this corpus), a change of -0.42. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.03, then +0.24, then +0.23 — a plateau around step 1, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores -0.21 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.09 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (-0.21, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 18 s · snippets
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.206 before conversion and 0.256 after — it rose by 0.462. Neighbour-to-neighbour the worst pair went -0.089 → 0.256. (The earlier render, with segment 1 left raw, scores 0.117 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.494 in the original and +0.137 after conversion — 28 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Intoxication Altered States of Consciousness, -0.426 became -0.498.
Quality. Mean predicted overall quality across the segments went 2.69 → 2.75 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 -0.206 → 0.256+0.462identity cos neighbours -0.089 → 0.256d_b rescored +0.494 → +0.137d_a rescored -0.426 → -0.498d_a mined -0.425d_b mined 0.494min_cos_consec (site) -0.0856min_cos_anchor (site) -0.2058dataset snippetslang ?speaker batch247_part4_batch247_patotal 16.9schain gain +1.0 dBseam step 2.3 dBcrossfades 150/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, balanced body, measured, normally alert, fairly steady
(relaxed, frequent disfluency, slurred, casual)मैं जो बोता हूं वह तीन से चार एकड़ गेहूं में बोता हूं।
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 2.9/10; 4.9s.
batch247_part4_batch247_part4_chunk_693_1_496782 · in -30.7 dBFS · gain +10.7 dB · snippets-00771
(slightly relaxed, no disfluency, clear, formal)around 7 million tons were exported to other countries.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.6/10; 3.5s.
batch247_part4_batch247_part4_chunk_693_1_496793 · in -29.7 dBFS · gain +9.7 dB · snippets-00771
(emotional numbness· slightly relaxed, little disfluency, clear, formal)After 107 million tons of wheat produced in 2021,
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as emotional numbness; style: formal, monologue; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.5/10; 4.5s.
batch247_part4_batch247_part4_chunk_693_1_496817 · in -28.0 dBFS · gain +8.0 dB · snippets-00771
(slightly relaxed, no disfluency, clear, narration)India ranked 101 out of 116 countries.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: narration, storytelling; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.1/10; 4.4s.
batch247_part4_batch247_part4_chunk_693_1_497240 · in -30.8 dBFS · gain +10.8 dB · snippets-00771
This chain comes from the proxy rule: the same two-sided test as above, but because Fear is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Fear around average — 0.56, higher than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.40.
At the same time Emotional Numbness goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.32 (lower than 68 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.01, then +0.19 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores -0.14 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.07 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (-0.14, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 25 s · snippets
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.123 before conversion and 0.503 after — it rose by 0.626. Neighbour-to-neighbour the worst pair went -0.045 → 0.455. (The earlier render, with segment 1 left raw, scores 0.473 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.400 in the original and +0.576 after conversion — 144 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, -0.394 became -0.605.
Quality. Mean predicted overall quality across the segments went 2.60 → 2.75 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 -0.123 → 0.503+0.626identity cos neighbours -0.045 → 0.455d_b rescored +0.400 → +0.576d_a rescored -0.394 → -0.605d_a mined -0.396d_b mined 0.400min_cos_consec (site) -0.0665min_cos_anchor (site) -0.1356dataset snippetslang ?speaker batch77_part1_batch77_parttotal 24.1schain gain +1.5 dBseam step 1.4 dBcrossfades 150/100/100 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a child masculine voice · measured
(subdued, relaxed, steady, monologue)is (low mumble) uh, meet the sea. So, due to sea level rise, (low mumble) um, there's a large number of sun energy infusion which is
full caption & clip details
A child masculine voice; delivery is subdued, measured, relaxed, steady; timbre is slightly cool, dark, slightly rough, slightly thin; slurred, frequent disfluency, narrow pitch range, audible breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: monologue, casual; below-average recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.6/10; 8.4s.
batch77_part1_batch77_part1_chunk_1690_1_1998559 · in -24.4 dBFS · gain +4.4 dB · snippets-01284
(doubt, longing, distress·normally alert, slightly relaxed, fairly steady, casual)a little bit about Bangladesh government right now, but how is the relationship going with them?
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as doubt, longing, distress; style: casual, monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 2.5/10; 5.1s.
batch77_part1_batch77_part1_chunk_1690_1_1998641 · in -23.9 dBFS · gain +3.9 dB · snippets-01284
(subdued, slightly relaxed, fairly steady, casual)policies that has been amended in terms of including climate change
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; below-average recording, quiet background; genuineness 3.4/6; vocal-burst blend 2.6/10; 5.1s.
batch77_part1_batch77_part1_chunk_1690_1_1998721 · in -25.9 dBFS · gain +5.9 dB · snippets-01284
(fear·normally alert, slightly relaxed, fairly steady, whispered)just like systems or things that can change that would help you in your battle against climate change.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as fear; style: whispered, monologue; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.4/10; 6.0s.
batch77_part1_batch77_part1_chunk_1690_1_1998736 · in -25.4 dBFS · gain +5.5 dB · snippets-01284
This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Infatuation clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.28.
At the same time Astonishment Surprise goes the other way, from 0.77 (higher than 77 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are -0.09, then +0.16, then +0.21 — not a clean run: step 1 moves back the other way by 0.09 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.70 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.70 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.70, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 23 s · snippets
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.705 before conversion and 0.811 after — it rose by 0.107. Neighbour-to-neighbour the worst pair went 0.705 → 0.811. (The earlier render, with segment 1 left raw, scores 0.720 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.278 in the original and +0.182 after conversion — 65 % of the delta retained. On the other named axis, Astonishment Surprise, -0.407 became -0.591.
Quality. Mean predicted overall quality across the segments went 2.85 → 2.90 (+0.05) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.705 → 0.811+0.107identity cos neighbours 0.705 → 0.811d_b rescored +0.278 → +0.182d_a rescored -0.407 → -0.591d_a mined -0.410d_b mined 0.280min_cos_consec (site) 0.6970min_cos_anchor (site) 0.6970dataset snippetslang ?speaker batch183_part2_batch183_patotal 22.4schain gain +2.4 dBseam step 0.9 dBcrossfades 100/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, fairly steady
(normal-paced, no disfluency, formal, monologue)is apparently derived from the name of the goddess Mumba.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 2.0/10; 3.1s.
batch183_part2_batch183_part2_chunk_263_1_402087 · in -25.9 dBFS · gain +6.0 dB · snippets-00434
(normal-paced, almost no disfluency, newsreading, formal)These cultural characteristics are defined by various influences, including geographical, spiritual, and agricultural considerations.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 8.4s.
batch183_part2_batch183_part2_chunk_263_1_402172 · in -26.4 dBFS · gain +6.4 dB · snippets-00434
(normal-paced, almost no disfluency, monologue, storytelling)His holiness, the Dalai Lama, is the highest political as well as spiritual authority of Tibet.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: monologue, storytelling; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 1.7/10; 5.7s.
batch183_part2_batch183_part2_chunk_263_1_402247 · in -25.3 dBFS · gain +5.3 dB · snippets-00434
(infatuation·brisk, almost no disfluency, conversational, narration)Do you find it difficult to make up your mind? A private fashion show with professional models is here to help you.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as infatuation; style: conversational, narration; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 2.2/10; 5.7s.
batch183_part2_batch183_part2_chunk_263_1_402302 · in -25.9 dBFS · gain +5.9 dB · snippets-00434
Fatigue Exhaustion ↓ / Intoxication Altered States of Consciousness ↑identity +0.31emotion 51 % c-snippets-PXR · #18
This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Intoxication Altered States of Consciousness around average — 0.53, higher than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.40.
At the same time Fatigue Exhaustion goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.57 (higher than 57 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.24 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.04 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.02 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.04, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 13 s · snippets
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.018 before conversion and 0.330 after — it rose by 0.313. Neighbour-to-neighbour the worst pair went -0.006 → 0.415. (The earlier render, with segment 1 left raw, scores 0.174 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.404 in the original and +0.207 after conversion — 51 % of the delta retained. On the other named axis, Fatigue Exhaustion, -0.328 became -0.417.
Quality. Mean predicted overall quality across the segments went 2.63 → 2.75 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.018 → 0.330+0.313identity cos neighbours -0.006 → 0.415d_b rescored +0.404 → +0.207d_a rescored -0.328 → -0.417d_a mined -0.328d_b mined 0.403min_cos_consec (site) 0.0171min_cos_anchor (site) 0.0411dataset snippetslang ?speaker batch212_part0_batch212_patotal 12.5schain gain +2.8 dBseam step 3.8 dBcrossfades 100/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, normally alert, slightly relaxed, light breath
(measured, fairly steady, some disfluency, whispered)We definitely started to see a lot of traction coming out of the Philippines, especially after
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: whispered, monologue; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 3.4/10; 5.8s.
batch212_part0_batch212_part0_chunk_379_1_37395 · in -28.8 dBFS · gain +8.8 dB · snippets-00583
(contemplation· measured, steady, no disfluency, monologue)fundamental change in the nature of work.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, dark, fairly smooth, balanced body; somewhat unclear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation; style: monologue, whispered; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 3.8/10; 3.0s.
batch212_part0_batch212_part0_chunk_379_1_37418 · in -29.8 dBFS · gain +9.8 dB · snippets-00583
(intoxication altered states of consciousness·fast, fairly steady, frequent disfluency, casual)Kahit po sana kumita po ng kahit 3000 po a month.
full caption & clip details
A child feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, thin; slurred, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 5.1/10; 4.0s.
batch212_part0_batch212_part0_chunk_379_1_37532 · in -15.8 dBFS · gain -4.2 dB · snippets-00583
This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Emotional Numbness clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.30.
At the same time Impatience and Irritability goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.10, then +0.21, then -0.02 — not a clean run: step 3 moves back the other way by 0.02 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.06 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.06 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.06, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 16 s · snippets
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.072 before conversion and 0.329 after — it rose by 0.257. Neighbour-to-neighbour the worst pair went 0.065 → 0.286. (The earlier render, with segment 1 left raw, scores 0.066 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.297 in the original and +0.152 after conversion — 51 % of the delta retained. On the other named axis, Impatience and Irritability, -0.329 became -0.314.
Quality. Mean predicted overall quality across the segments went 2.60 → 2.75 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.072 → 0.329+0.257identity cos neighbours 0.065 → 0.286d_b rescored +0.297 → +0.152d_a rescored -0.329 → -0.314d_a mined -0.327d_b mined 0.300min_cos_consec (site) 0.0597min_cos_anchor (site) 0.0610dataset snippetslang ?speaker batch112_part3_batch112_patotal 15.2schain gain +0.5 dBseam step 1.6 dBcrossfades 150/100/100 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · slightly relaxed, fairly steady
(impatience and irritability, intoxication altered states of consciousness, distress · fast, normally alert, some disfluency, casual)訳達者、その芝居出しとったところが手前から約
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is slightly cool, dark, very rough, thin; slurred, some disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, guarded; reads as impatience and irritability, intoxication altered states of consciousness, distress; style: casual, playful; poor recording, quiet background; genuineness 4.6/6; vocal-burst blend 4.6/10; 3.4s.
batch112_part3_batch112_part3_chunk_2007_1_2065547 · in -32.2 dBFS · gain +12.2 dB · snippets-00066
(pain· fast, normally alert, some disfluency, casual)(low mumble) 杨教授呢,就是因为杨教授是担任这个章程的翻译。
full caption & clip details
A child feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, thin; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 5.5/10; 3.9s.
batch112_part3_batch112_part3_chunk_2007_1_2065634 · in -33.4 dBFS · gain +13.4 dB · snippets-00066
(sadness, pain, helplessness·measured, subdued, no disfluency, whispered)very cognizant of losing the person she once was.
full caption & clip details
An elderly feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is slightly warm, dark, fairly smooth, thin; clear, no disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, neutral openness; reads as sadness, pain, helplessness; style: whispered, monologue; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 3.2/10; 3.7s.
batch112_part3_batch112_part3_chunk_2007_1_2065646 · in -35.6 dBFS · gain +15.6 dB · snippets-00066
(emotional numbness·normal-paced, normally alert, little disfluency, casual)One person, actually just one idea, can start a war or end one.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as emotional numbness; style: casual, storytelling; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 0.5/10; 4.6s.
batch112_part3_batch112_part3_chunk_2007_1_2065721 · in -35.1 dBFS · gain +15.1 dB · snippets-00066
This chain comes from the proxy rule: the same two-sided test as above, but because Astonishment Surprise is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Astonishment Surprise around average — 0.52, higher than 52 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.35.
At the same time Pride goes the other way, from 0.88 (higher than 88 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.35. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.13 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.10 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.10 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.10, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 14 s · snippets
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.033 before conversion and 0.341 after — it rose by 0.307. Neighbour-to-neighbour the worst pair went 0.033 → 0.341. (The earlier render, with segment 1 left raw, scores 0.199 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.347 in the original and +0.582 after conversion — 167 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Pride, -0.354 became -0.053.
Quality. Mean predicted overall quality across the segments went 2.66 → 2.88 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.033 → 0.341+0.307identity cos neighbours 0.033 → 0.341d_b rescored +0.347 → +0.582d_a rescored -0.354 → -0.053d_a mined -0.355d_b mined 0.348min_cos_consec (site) 0.1011min_cos_anchor (site) 0.1011dataset snippetslang ?speaker batch45_part2_batch45_parttotal 13.8schain gain +2.0 dBseam step 1.6 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-bright, normal-paced, light breath
(normally alert, slightly relaxed, fairly steady, formal)his hard man reputation and formidable record.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.9/10; 3.0s.
batch45_part2_batch45_part2_chunk_1404_1_1787228 · in -22.3 dBFS · gain +2.3 dB · snippets-01122
(helplessness, disappointment, confusion·energised, neutral tension, moderately variable, casual)Nobody wanted to give it to me, so I just went by my head and did what I thought was the best thing for me to do.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as helplessness, disappointment, confusion; style: casual, storytelling; below-average recording, quiet background; genuineness 3.3/6; vocal-burst blend 3.5/10; 6.2s.
batch45_part2_batch45_part2_chunk_1404_1_1787270 · in -22.1 dBFS · gain +2.1 dB · snippets-01122
(normally alert, slightly relaxed, fairly steady, narration)from the first bell, the crowd expected Liston to march over and knock Clay out.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 1.0/10; 4.9s.
batch45_part2_batch45_part2_chunk_1404_1_1787296 · in -22.0 dBFS · gain +2.0 dB · snippets-01122