c-snippets-AB2 — voice-corrected

Corpus snippets in isolation, rule AB2.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_c-snippets-AB2.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
51segments re-voiced
0.162 → 0.565median worst-to-anchor identity cosine
72 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Hope Enthusiasm Optimism ↓  /  Confusionidentity +0.43 emotion 35 %   c-snippets-AB2 · #1

This chain comes from the two-sided rule: it only counts if both emotions move — Hope Enthusiasm Optimism down and Confusion up — by at least 0.25 each.

The chain starts with Confusion below average — 0.34, lower than 66 % of clips in this corpus — and ends with it strongly present at 0.79, higher than 79 % of clips in this corpus. That is a total rise of 0.45.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.14 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.15 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.14, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 19 s · snippets

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.149 before conversion and 0.576 after — it rose by 0.427. Neighbour-to-neighbour the worst pair went 0.155 → 0.700. (The earlier render, with segment 1 left raw, scores 0.447 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.451 in the original and +0.157 after conversion — 35 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Hope Enthusiasm Optimism, -0.268 became -0.363.

Quality. Mean predicted overall quality across the segments went 2.82 → 2.95 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.149 → 0.576 +0.427identity cos neighbours 0.155 → 0.700d_b rescored +0.451 → +0.157d_a rescored -0.268 → -0.363d_a mined -0.270d_b mined 0.451min_cos_consec (site) 0.1530min_cos_anchor (site) 0.1357dataset snippetslang ?speaker batch56_part1_batch56_parttotal 18.6schain gain +0.6 dBseam step 4.6 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, good recording, normal-paced, normally alert, average clarity, light breath
(hope enthusiasm optimism · slightly relaxed, fairly steady, some disfluency, casual) that they will be, they think they'll be happy for the rest of their lives within three months.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as hope enthusiasm optimism; style: casual, conversational; good recording, no background noise; genuineness 2.9/6; vocal-burst blend 3.1/10; 4.6s.
batch56_part1_batch56_part1_chunk_1502_1_1996029 · in -21.7 dBFS · gain +1.7 dB · snippets-01175
(thankfulness gratitude, affection, pleasure ecstasy · neutral tension, fairly steady, some disfluency, conversational) there is actually scientific evidence showing that when I appreciate my my partner, when I appreciate my children, when I appreciate my work.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as thankfulness gratitude, affection, pleasure ecstasy; style: conversational, casual; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 7.4/10; 7.2s.
batch56_part1_batch56_part1_chunk_1502_1_1996165 · in -22.5 dBFS · gain +2.5 dB · snippets-01175
(slightly relaxed, moderately variable, frequent disfluency, casual) And (ahem) um, I would like you to explain briefly what's that MPA's process about, and (ahem) um, yeah.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 2.9/6; vocal-burst blend 1.2/10; 7.1s.
batch56_part1_batch56_part1_chunk_1502_1_1996212 · in -24.1 dBFS · gain +4.1 dB · snippets-01175
Astonishment Surprise ↓  /  Intoxication Altered States of Consciousnessidentity +0.41 emotion 85 %   c-snippets-AB2 · #2

This chain comes from the two-sided rule: it only counts if both emotions move — Astonishment Surprise down and Intoxication Altered States of Consciousness up — by at least 0.25 each.

The chain starts with Intoxication Altered States of Consciousness clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.25.

At the same time Astonishment Surprise goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.07, then +0.05, then +0.13 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.14 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.08 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.14, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 14 s · snippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.144 before conversion and 0.554 after — it rose by 0.410. Neighbour-to-neighbour the worst pair went 0.066 → 0.643. (The earlier render, with segment 1 left raw, scores 0.397 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.251 in the original and +0.213 after conversion — 85 % of the delta retained, which is most of it. On the other named axis, Astonishment Surprise, -0.336 became -0.338.

Quality. Mean predicted overall quality across the segments went 2.46 → 2.60 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.144 → 0.554 +0.410identity cos neighbours 0.066 → 0.643d_b rescored +0.251 → +0.213d_a rescored -0.336 → -0.338d_a mined -0.336d_b mined 0.251min_cos_consec (site) 0.0829min_cos_anchor (site) 0.1401dataset snippetslang ?speaker batch166_part3_batch166_patotal 12.9schain gain +2.5 dBseam step 1.0 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, normally alert
(astonishment surprise, confusion, doubt · normal-paced, slightly relaxed, moderately variable, casual) What? Yeah, that, that doesn't mean anything.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as astonishment surprise, confusion, doubt; style: casual, conversational; average recording, no background noise; genuineness 4.5/6; vocal-burst blend 1.0/10; 3.2s.
batch166_part3_batch166_part3_chunk_2491_1_2644822 · in -22.7 dBFS · gain +2.7 dB · snippets-00351
(amusement, teasing · normal-paced, slightly relaxed, moderately variable, casual) Okay. (surprised gasp) So I was like, I'm gonna go for it, and he went.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as amusement, teasing; style: casual, conversational; good recording, no background noise; genuineness 4.6/6; vocal-burst blend 5.1/10; 3.0s.
batch166_part3_batch166_part3_chunk_2491_1_2644880 · in -21.3 dBFS · gain +1.3 dB · snippets-00351
(longing, infatuation, relief · normal-paced, slightly relaxed, fairly steady, casual) on like that all that day leading up to my special, and we took
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as longing, infatuation, relief; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 5.3/10; 4.2s.
batch166_part3_batch166_part3_chunk_2491_1_2644900 · in -24.3 dBFS · gain +4.3 dB · snippets-00351
(intoxication altered states of consciousness, confusion · slow, fully relaxed, fairly steady, casual) in green room and the clumps come in.
full caption & clip details
A young adult masculine voice; delivery is normally alert, slow, fully relaxed, fairly steady; timbre is neutral-toned, dark, slightly rough, thin; slurred, frequent disfluency, moderate pitch range, audible breath; affect is mildly negative, neutral stance, neutral openness; reads as intoxication altered states of consciousness, confusion; style: casual, monologue; average recording, quiet background; explicit content; genuineness 4.0/6; vocal-burst blend 1.9/10; 3.2s.
batch166_part3_batch166_part3_chunk_2491_1_2644991 · in -27.7 dBFS · gain +7.7 dB · snippets-00351
Pride ↓  /  Contemplationidentity +0.66 emotion 50 %   c-snippets-AB2 · #3

This chain comes from the two-sided rule: it only counts if both emotions move — Pride down and Contemplation up — by at least 0.25 each.

The chain starts with Contemplation clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.25.

At the same time Pride goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.37. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.00 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.06 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.06 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.06, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 32 s · snippets

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.064 before conversion and 0.592 after — it rose by 0.656. Neighbour-to-neighbour the worst pair went -0.064 → 0.537. (The earlier render, with segment 1 left raw, scores 0.494 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.247 in the original and +0.123 after conversion — 50 % of the delta retained. On the other named axis, Pride, -0.367 became -0.544.

Quality. Mean predicted overall quality across the segments went 2.96 → 3.12 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 -0.064 → 0.592 +0.656identity cos neighbours -0.064 → 0.537d_b rescored +0.247 → +0.123d_a rescored -0.367 → -0.544d_a mined -0.366d_b mined 0.254min_cos_consec (site) -0.0636min_cos_anchor (site) -0.0636dataset snippetslang ?speaker batch131_part3_batch131_patotal 31.8schain gain +2.6 dBseam step 0.4 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, no background noise, slightly relaxed
(measured, normally alert, fairly steady, monologue) A lot of times doing the work of creating a formula or writing a book makes you better at what you do.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.2/10; 9.0s.
batch131_part3_batch131_part3_chunk_2178_1_2400246 · in -20.3 dBFS · gain +0.3 dB · snippets-00168
(longing, triumph, relief · slow, very low-energy, moderately variable, narration) But it melted away as fast, for we made a lane through us for a single ray from the fire to fall on the face of the little sprite. And he thought it was a child of his own that had died when just the age of her child niece.
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, slow, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, slightly thin; clear, almost no disfluency, very wide pitch range, no audible breath; affect is mildly negative, neutral stance, neutral openness; reads as longing, triumph, relief; style: narration, whispered; average recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 17.0s.
batch131_part3_batch131_part3_chunk_2178_1_2400270 · in -29.4 dBFS · gain +9.4 dB · snippets-00168
(contemplation, fear · normal-paced, normally alert, fairly steady, ASMR) I think there are a lot of people like us though, in this country, and especially in, I don't know, what we call the Bible Belt.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contemplation, fear; style: ASMR, conversational; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 3.9/10; 6.1s.
batch131_part3_batch131_part3_chunk_2178_1_2400329 · in -24.5 dBFS · gain +4.5 dB · snippets-00168
Fatigue Exhaustion ↓  /  Reliefidentity +0.44 emotion 136 %   c-snippets-AB2 · #4

This chain comes from the two-sided rule: it only counts if both emotions move — Fatigue Exhaustion down and Relief up — by at least 0.25 each.

The chain starts with Relief around average — 0.50, right about the corpus median — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.47.

At the same time Fatigue Exhaustion goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.50 (right about the corpus median), a change of -0.43. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.19, then +0.23, then +0.05 — a plateau around step 3, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.20 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.14 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.20, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 22 s · snippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.180 before conversion and 0.622 after — it rose by 0.442. Neighbour-to-neighbour the worst pair went 0.071 → 0.499. (The earlier render, with segment 1 left raw, scores 0.400 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.472 in the original and +0.641 after conversion — 136 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Fatigue Exhaustion, -0.435 became -0.159.

Quality. Mean predicted overall quality across the segments went 2.21 → 2.76 (+0.55) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.180 → 0.622 +0.442identity cos neighbours 0.071 → 0.499d_b rescored +0.472 → +0.641d_a rescored -0.435 → -0.159d_a mined -0.434d_b mined 0.470min_cos_consec (site) 0.1432min_cos_anchor (site) 0.1976dataset snippetslang ?speaker batch211_part2_batch211_patotal 21.6schain gain +0.5 dBseam step 1.0 dBcrossfades 100/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth
(fatigue exhaustion, longing · slow, subdued, slightly relaxed, whispered) will also improve the time and quality of items stemming from factory.
full caption & clip details
A young adult masculine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, longing; style: whispered, monologue; average recording, no background noise; genuineness 1.7/6; vocal-burst blend 1.4/10; 4.9s.
batch211_part2_batch211_part2_chunk_369_1_582055 · in -15.9 dBFS · gain -4.1 dB · snippets-00580
(confusion, doubt, fatigue exhaustion · slow, very low-energy, relaxed, casual) there would be less like injuries or less like
full caption & clip details
A child feminine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is neutral-toned, dark, fairly smooth, slightly thin; slurred, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as confusion, doubt, fatigue exhaustion; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.1/6; vocal-burst blend 6.9/10; 4.5s.
batch211_part2_batch211_part2_chunk_369_1_582107 · in -19.4 dBFS · gain -0.6 dB · snippets-00580
(hope enthusiasm optimism, contentment, contemplation · normal-paced, normally alert, slightly relaxed, casual) some of the benefits would be that business will be able to communicate together like better as well as like everything will be more efficient and effective.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism, contentment, contemplation; style: casual, monologue; below-average recording, quiet background; genuineness 4.6/6; vocal-burst blend 6.6/10; 9.6s.
batch211_part2_batch211_part2_chunk_369_1_582391 · in -20.6 dBFS · gain +0.6 dB · snippets-00580
(relief · measured, normally alert, slightly relaxed, casual) We believe this information may help but also
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as relief; style: casual, conversational; average recording, no background noise; genuineness 3.9/6; vocal-burst blend 3.8/10; 3.0s.
batch211_part2_batch211_part2_chunk_369_1_582436 · in -13.9 dBFS · gain -6.1 dB · snippets-00580
Fear ↓  /  Sexual Lustidentity +0.07 emotion 36 %   c-snippets-AB2 · #5

This chain comes from the two-sided rule: it only counts if both emotions move — Fear down and Sexual Lust up — by at least 0.25 each.

The chain starts with Sexual Lust around average — 0.44, lower than 56 % of clips in this corpus — and ends with it clearly present at 0.70, higher than 70 % of clips in this corpus. That is a total rise of 0.26.

At the same time Fear goes the other way, from 0.82 (higher than 82 % of clips in this corpus) to 0.49 (right about the corpus median), a change of -0.32. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.09, then +0.17 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.68 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.67 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.68, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 13 s · snippets

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.643 before conversion and 0.713 after — it rose by 0.070. Neighbour-to-neighbour the worst pair went 0.667 → 0.713. (The earlier render, with segment 1 left raw, scores 0.551 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sexual Lust moved +0.265 in the original and +0.096 after conversion — 36 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Fear, -0.344 became +0.083.

Quality. Mean predicted overall quality across the segments went 2.40 → 2.65 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.643 → 0.713 +0.070identity cos neighbours 0.667 → 0.713d_b rescored +0.265 → +0.096d_a rescored -0.344 → +0.083d_a mined -0.323d_b mined 0.262min_cos_consec (site) 0.6740min_cos_anchor (site) 0.6835dataset snippetslang ?speaker batch33_part4_batch33_parttotal 12.6schain gain +2.3 dBseam step 1.3 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, slightly dark, no background noise, normally alert, slightly relaxed, fairly steady, moderate pitch range, light breath
(measured, almost no disfluency, average clarity, formal) Note that the casting unit starts operating automatically as each matrix line is delivered.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 0.1/10; 5.8s.
batch33_part4_batch33_part4_chunk_1292_1_1015922 · in -24.4 dBFS · gain +4.3 dB · snippets-01065
(measured, almost no disfluency, average clarity, formal) The keyboard has 90 keys on six rows of 15 each.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; average recording, no background noise; genuineness 2.1/6; vocal-burst blend 0.6/10; 3.8s.
batch33_part4_batch33_part4_chunk_1292_1_1015954 · in -25.9 dBFS · gain +5.9 dB · snippets-01065
(normal-paced, some disfluency, somewhat unclear, casual) which turns them in correct sequence to the assembling elevator.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, no background noise; genuineness 3.2/6; vocal-burst blend 1.2/10; 3.4s.
batch33_part4_batch33_part4_chunk_1292_1_1016006 · in -27.0 dBFS · gain +7.0 dB · snippets-01065
Pain ↓  /  Intoxication Altered States of Consciousnessidentity +0.39 emotion REVERSED   c-snippets-AB2 · #6

This chain comes from the two-sided rule: it only counts if both emotions move — Pain down and Intoxication Altered States of Consciousness up — by at least 0.25 each.

The chain starts with Intoxication Altered States of Consciousness around average — 0.47, lower than 53 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.37.

At the same time Pain goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.37. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.18 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.17 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.19 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.17, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 17 s · snippets

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.148 before conversion and 0.539 after — it rose by 0.391. Neighbour-to-neighbour the worst pair went 0.168 → 0.450. (The earlier render, with segment 1 left raw, scores 0.484 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

The emotional move did not survive. Re-scored end to end, Intoxication Altered States of Consciousness moved +0.378 in the original and -0.126 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Pain, -0.371 became -0.193.

Quality. Mean predicted overall quality across the segments went 2.70 → 2.79 (+0.09) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.148 → 0.539 +0.391identity cos neighbours 0.168 → 0.450d_b rescored +0.378 → -0.126d_a rescored -0.371 → -0.193d_a mined -0.371d_b mined 0.374min_cos_consec (site) 0.1874min_cos_anchor (site) 0.1666dataset snippetslang ?speaker batch265_part2_batch265_patotal 16.2schain gain +1.5 dBseam step 2.4 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, fairly smooth, normally alert, fairly steady, moderate pitch range
(pain · normal-paced, slightly relaxed, no disfluency, formal) In winter of 2015, Paul made national news when he took to the street with a shovel to clear icy sidewalks.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as pain; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.2/10; 7.1s.
batch265_part2_batch265_part2_chunk_851_1_738440 · in -25.7 dBFS · gain +5.7 dB · snippets-00860
(measured, slightly relaxed, no disfluency, formal) friend and executive director of the Nova Scotia Accessibility Directorate, Jerry Post.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.4/10; 5.4s.
batch265_part2_batch265_part2_chunk_851_1_738481 · in -26.5 dBFS · gain +6.5 dB · snippets-00860
(measured, fully relaxed, frequent disfluency, casual) you can just pick anything up and and emulate it or play it well.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, fully relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, thin; slurred, frequent disfluency, moderate pitch range, audible breath; affect is mildly negative, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 4.6/10; 4.0s.
batch265_part2_batch265_part2_chunk_851_1_738503 · in -31.9 dBFS · gain +11.9 dB · snippets-00860
Relief ↓  /  Emotional Numbnessidentity +0.48 emotion 152 %   c-snippets-AB2 · #7

This chain comes from the two-sided rule: it only counts if both emotions move — Relief down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.30.

At the same time Relief goes the other way, from 0.80 (higher than 80 % of clips in this corpus) to 0.47 (lower than 53 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.12, then +0.19 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.00 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.00 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.00, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 11 s · snippets

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.007 before conversion and 0.470 after — it rose by 0.478. Neighbour-to-neighbour the worst pair went -0.007 → 0.470. (The earlier render, with segment 1 left raw, scores 0.401 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.306 in the original and +0.465 after conversion — 152 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Relief, -0.369 became -0.360.

Quality. Mean predicted overall quality across the segments went 2.72 → 2.79 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 -0.007 → 0.470 +0.478identity cos neighbours -0.007 → 0.470d_b rescored +0.306 → +0.465d_a rescored -0.369 → -0.360d_a mined -0.333d_b mined 0.305min_cos_consec (site) -0.0040min_cos_anchor (site) -0.0040dataset snippetslang ?speaker batch46_part3_batch46_parttotal 10.3schain gain +1.7 dBseam step 1.3 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, measured, slightly relaxed
(normally alert, formal, monologue) Surf Shark VPN may help you to deal with that.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 4.5/10; 3.4s.
batch46_part3_batch46_part3_chunk_1408_1_1848551 · in -16.8 dBFS · gain -3.2 dB · snippets-01128
(awe, affection, infatuation · very low-energy, monologue, whispered) it creates a deep emotional connection that transcends words.
full caption & clip details
An adult feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as awe, affection, infatuation; style: monologue, whispered; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 4.3/10; 4.2s.
batch46_part3_batch46_part3_chunk_1408_1_1848756 · in -26.0 dBFS · gain +6.0 dB · snippets-01128
(emotional numbness, fear, distress · normally alert, monologue, storytelling) a decrease in heart rate and shallower breathing.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, fear, distress; style: monologue, storytelling; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 3.5/10; 3.1s.
batch46_part3_batch46_part3_chunk_1408_1_1848813 · in -27.3 dBFS · gain +7.3 dB · snippets-01128
Malevolence Malice ↓  /  Contemplationidentity −0.02 emotion 62 %   c-snippets-AB2 · #8

This chain comes from the two-sided rule: it only counts if both emotions move — Malevolence Malice down and Contemplation up — by at least 0.25 each.

The chain starts with Contemplation around average — 0.52, higher than 52 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.44.

At the same time Malevolence Malice goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.22, then +0.01, then +0.21 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.80 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.84 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.80, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 29 s · snippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.765 before conversion and 0.747 after — it fell by 0.018. Neighbour-to-neighbour the worst pair went 0.765 → 0.747. (The earlier render, with segment 1 left raw, scores 0.710 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.433 in the original and +0.267 after conversion — 62 % of the delta retained. On the other named axis, Malevolence Malice, -0.293 became -0.379.

Quality. Mean predicted overall quality across the segments went 2.99 → 3.04 (+0.05) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.765 → 0.747 -0.018identity cos neighbours 0.765 → 0.747d_b rescored +0.433 → +0.267d_a rescored -0.293 → -0.379d_a mined -0.294d_b mined 0.440min_cos_consec (site) 0.8382min_cos_anchor (site) 0.7984dataset snippetslang ?speaker batch232_part2_batch232_patotal 28.0schain gain +2.0 dBseam step 1.0 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, measured, normally alert
(malevolence malice, awe, shame · fairly steady, no disfluency, monologue, narration) And killed that sea creature known as a Nahiru.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, awe, shame; style: monologue, narration; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 4.3/10; 3.6s.
batch232_part2_batch232_part2_chunk_555_1_554824 · in -28.7 dBFS · gain +8.7 dB · snippets-00689
(pride, fear, malevolence malice · fairly steady, almost no disfluency, narration, formal) and instead, set about destroying the cities with the same viciousness that the Assyrians had once reserved for the cities of Elam.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, fear, malevolence malice; style: narration, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.8/10; 9.1s.
batch232_part2_batch232_part2_chunk_555_1_554836 · in -27.8 dBFS · gain +7.8 dB · snippets-00689
(sadness, disappointment · steady, almost no disfluency, narration, formal) For centuries now, the powerful Elamites had been their rivals and kept their ambitions in check.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sadness, disappointment; style: narration, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.3/10; 7.2s.
batch232_part2_batch232_part2_chunk_555_1_554870 · in -28.1 dBFS · gain +8.1 dB · snippets-00689
(contemplation, emotional numbness · steady, almost no disfluency, narration, monologue) and it shows that in the social upheaval of this period of chaos, some of the power of the nobility was being eroded.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, emotional numbness; style: narration, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.8/10; 8.7s.
batch232_part2_batch232_part2_chunk_555_1_554887 · in -27.6 dBFS · gain +7.6 dB · snippets-00689
Contempt ↓  /  Intoxication Altered States of Consciousnessidentity −0.06 emotion 71 %   c-snippets-AB2 · #9

This chain comes from the two-sided rule: it only counts if both emotions move — Contempt down and Intoxication Altered States of Consciousness up — by at least 0.25 each.

The chain starts with Intoxication Altered States of Consciousness clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.34.

At the same time Contempt goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.09, then +0.24 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.05 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.02 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.05, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 73 s · snippets

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.342 before conversion and 0.285 after — it fell by 0.057. Neighbour-to-neighbour the worst pair went 0.061 → 0.048. (The earlier render, with segment 1 left raw, scores 0.395 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.337 in the original and +0.240 after conversion — 71 % of the delta retained, which is most of it. On the other named axis, Contempt, -0.338 became -0.428.

Quality. Mean predicted overall quality across the segments went 2.88 → 3.09 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.342 → 0.285 -0.057identity cos neighbours 0.061 → 0.048d_b rescored +0.337 → +0.240d_a rescored -0.338 → -0.428d_a mined -0.337d_b mined 0.337min_cos_consec (site) 0.0216min_cos_anchor (site) -0.0521dataset snippetslang ?speaker batch88_part2_batch88_parttotal 71.9schain gain +2.4 dBseam step 0.4 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, quiet background, normally alert, light breath
(contempt, impatience and irritability, disgust · normal-paced, slightly relaxed, fairly steady, conversational) and saying he will stop looking and let her look for him. He would then go on to tweet on the 10th of April that there's some women he would do anything for.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contempt, impatience and irritability, disgust; style: conversational, formal; good recording, quiet background; genuineness 2.2/6; vocal-burst blend 2.3/10; 6.8s.
batch88_part2_batch88_part2_chunk_1791_1_1449175 · in -27.5 dBFS · gain +7.5 dB · snippets-01345
(infatuation, jealousy and envy, sadness · brisk, slightly relaxed, fairly steady, monologue) somebody better not try him in his suit. So naturally, you can imagine, Von ended up spending a significant amount of time in trouble with the law as a teenager, with large patches of his most formative years spent inside jail cells. Von went to jail for the first time around August 2010 at age 16 for armed robbery. Apparently robbing someone for their car at gunpoint, something that he would actually later refer to on his song Armed and Dangerous, saying that he was arrested on August 11th, two days after his 17th birthday when he was offered 21 to 45 years in a plea deal. Von would end up spending time in boot camp, an intense military style juvenile program, ultimately rewarding him with early release. But soon after that initial charge, Von would catch another in January 2011, ultimately spending 15 months behind bars as a juvenile, being locked up between December 2010 and March 2012. And during this time of incarceration, Von's best friend T-Roy would tweet calling for his release. Eventually, in March 2012, Von would be a free man once again, but he was only free for a matter of months, from March to November 2012. But during this time, Von would undergo a transformation, going from a gun-toting stick-up kid.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as infatuation, jealousy and envy, sadness; style: monologue, casual; good recording, quiet background; genuineness 0.6/6; vocal-burst blend 4.3/10; 60.0s.
batch88_part2_batch88_part2_chunk_1791_1_1449205 · in -26.5 dBFS · gain +6.5 dB · snippets-01345
(intoxication altered states of consciousness, confusion, helplessness · fast, neutral tension, moderately variable, casual) We calling, we calling, we calling. I'm like, what's going on? What's going on? You got hit, you got hit, you got hit once, some shit, woah.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, slightly rough, thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as intoxication altered states of consciousness, confusion, helplessness; style: casual, conversational; below-average recording, quiet background; genuineness 5.6/6; vocal-burst blend 9.9/10; 5.5s.
batch88_part2_batch88_part2_chunk_1791_1_1449342 · in -22.9 dBFS · gain +2.9 dB · snippets-01345
Pain ↓  /  Fatigue Exhaustionidentity +0.35 emotion 71 %   c-snippets-AB2 · #10

This chain comes from the two-sided rule: it only counts if both emotions move — Pain down and Fatigue Exhaustion up — by at least 0.25 each.

The chain starts with Fatigue Exhaustion below average — 0.31, lower than 69 % of clips in this corpus — and ends with it strongly present at 0.79, higher than 79 % of clips in this corpus. That is a total rise of 0.47.

At the same time Pain goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.35 (lower than 65 % of clips in this corpus), a change of -0.55. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.11, then +0.17, then +0.19 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.02 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.09 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.02, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 23 s · snippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.030 before conversion and 0.384 after — it rose by 0.354. Neighbour-to-neighbour the worst pair went -0.108 → 0.384. (The earlier render, with segment 1 left raw, scores 0.306 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.472 in the original and +0.334 after conversion — 71 % of the delta retained, which is most of it. On the other named axis, Pain, -0.353 became -0.449.

Quality. Mean predicted overall quality across the segments went 2.70 → 2.79 (+0.08) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.030 → 0.384 +0.354identity cos neighbours -0.108 → 0.384d_b rescored +0.472 → +0.334d_a rescored -0.353 → -0.449d_a mined -0.549d_b mined 0.472min_cos_consec (site) -0.0946min_cos_anchor (site) -0.0210dataset snippetslang ?speaker batch214_part4_batch214_patotal 22.0schain gain +3.9 dBseam step 1.6 dBcrossfades 100/100/100 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, normally alert, slightly relaxed, moderate pitch range, light breath
(normal-paced, fairly steady, no disfluency, narration) the actual revolution is taking place out of the public eye.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 3.2/10; 3.8s.
batch214_part4_batch214_part4_chunk_396_1_279091 · in -17.4 dBFS · gain -2.6 dB · snippets-00594
(normal-paced, fairly steady, some disfluency, casual) now nowadays uh (low mumble) you (low mumble) know all the guests and crew members want to experience uh (low mumble) you know
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 2.3/10; 5.6s.
batch214_part4_batch214_part4_chunk_396_1_279118 · in -16.2 dBFS · gain -3.8 dB · snippets-00594
(measured, steady, almost no disfluency, newsreading) The magic body control chassis with curve function causes this body work to tilt by 2.65 degrees towards the inside of the bend.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 9.6s.
batch214_part4_batch214_part4_chunk_396_1_279133 · in -20.1 dBFS · gain +0.1 dB · snippets-00594
(normal-paced, fairly steady, frequent disfluency, casual) So with this (ahem) connectivity uh (low mumble) technology.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, playful; good recording, no background noise; genuineness 3.6/6; vocal-burst blend 2.5/10; 3.4s.
batch214_part4_batch214_part4_chunk_396_1_279380 · in -20.8 dBFS · gain +0.8 dB · snippets-00594
Triumph ↓  /  Concentrationidentity −0.03 emotion 30 %   c-snippets-AB2 · #11

This chain comes from the two-sided rule: it only counts if both emotions move — Triumph down and Concentration up — by at least 0.25 each.

The chain starts with Concentration below average — 0.41, lower than 59 % of clips in this corpus — and ends with it clearly present at 0.68, higher than 68 % of clips in this corpus. That is a total rise of 0.28.

At the same time Triumph goes the other way, from 0.84 (higher than 84 % of clips in this corpus) to 0.45 (lower than 55 % of clips in this corpus), a change of -0.39. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.04, then +0.24 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.74 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 12 s · snippets

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.810 before conversion and 0.781 after — it fell by 0.029. Neighbour-to-neighbour the worst pair went 0.725 → 0.781. (The earlier render, with segment 1 left raw, scores 0.672 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.275 in the original and +0.083 after conversion — 30 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Triumph, -0.389 became +0.219.

Quality. Mean predicted overall quality across the segments went 2.58 → 2.73 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.810 → 0.781 -0.029identity cos neighbours 0.725 → 0.781d_b rescored +0.275 → +0.083d_a rescored -0.389 → +0.219d_a mined -0.390d_b mined 0.276min_cos_consec (site) 0.7373min_cos_anchor (site) 0.8088dataset snippetslang ?speaker batch36_part0_batch36_parttotal 11.7schain gain +2.0 dBseam step 0.9 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(brisk, no disfluency, average clarity, casual) Rocksteady Studios released their fourth installment to the Arkham franchise.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, playful; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 2.1/10; 3.8s.
batch36_part0_batch36_part0_chunk_1320_1_746358 · in -24.6 dBFS · gain +4.6 dB · snippets-01075
(emotional numbness · normal-paced, no disfluency, average clarity, casual) each with their own architectural style and thematic element.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as emotional numbness; style: casual, conversational; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 4.0/10; 3.2s.
batch36_part0_batch36_part0_chunk_1320_1_746466 · in -24.5 dBFS · gain +4.5 dB · snippets-01075
(normal-paced, almost no disfluency, clear, casual) Alternatively, they can use the line launcher to create a zip line parallel to the ground.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, playful; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.8/10; 5.1s.
batch36_part0_batch36_part0_chunk_1320_1_746479 · in -24.4 dBFS · gain +4.4 dB · snippets-01075
Longing ↓  /  Malevolence Maliceidentity +0.49 emotion 111 %   c-snippets-AB2 · #12

This chain comes from the two-sided rule: it only counts if both emotions move — Longing down and Malevolence Malice up — by at least 0.25 each.

The chain starts with Malevolence Malice below average — 0.36, lower than 64 % of clips in this corpus — and ends with it strongly present at 0.75, higher than 75 % of clips in this corpus. That is a total rise of 0.39.

At the same time Longing goes the other way, from 0.99 (virtually no clip in this corpus scores higher) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.10 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.10 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.10, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 16 s · snippets

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.108 before conversion and 0.602 after — it rose by 0.494. Neighbour-to-neighbour the worst pair went 0.108 → 0.578. (The earlier render, with segment 1 left raw, scores 0.485 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Malevolence Malice moved +0.371 in the original and +0.413 after conversion — 111 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Longing, -0.335 became -0.388.

Quality. Mean predicted overall quality across the segments went 2.67 → 2.85 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.108 → 0.602 +0.494identity cos neighbours 0.108 → 0.578d_b rescored +0.371 → +0.413d_a rescored -0.335 → -0.388d_a mined -0.331d_b mined 0.389min_cos_consec (site) 0.1042min_cos_anchor (site) 0.1042dataset snippetslang ?speaker batch256_part1_batch256_patotal 14.9schain gain +1.1 dBseam step 1.3 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, balanced body, no background noise, normally alert, slightly relaxed, fairly steady, moderate pitch range, light breath
(longing · measured, some disfluency, average clarity, casual) At that time it was solid woods here, there was no lake in and we had to walk through a little.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as longing; style: casual, monologue; good recording, no background noise; genuineness 3.1/6; vocal-burst blend 1.7/10; 4.6s.
batch256_part1_batch256_part1_chunk_772_1_963149 · in -32.2 dBFS · gain +12.2 dB · snippets-00814
(contentment, thankfulness gratitude, affection · normal-paced, little disfluency, average clarity, whispered) you know, for myself as well, but especially for them, to see what it takes to build a house from the ground up.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contentment, thankfulness gratitude, affection; style: whispered, casual; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 5.4/10; 5.3s.
batch256_part1_batch256_part1_chunk_772_1_963178 · in -33.5 dBFS · gain +13.5 dB · snippets-00814
(measured, some disfluency, somewhat unclear, casual) Nothing freezes, but it keeps about the same temperature between 38 and 42.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; average recording, no background noise; genuineness 2.8/6; vocal-burst blend 2.0/10; 5.4s.
batch256_part1_batch256_part1_chunk_772_1_963220 · in -33.3 dBFS · gain +13.3 dB · snippets-00814
Hope Enthusiasm Optimism ↓  /  Emotional Numbnessidentity +0.28 emotion 299 %   c-snippets-AB2 · #13

This chain comes from the two-sided rule: it only counts if both emotions move — Hope Enthusiasm Optimism down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness around average — 0.54, higher than 54 % of clips in this corpus — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.29.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.59 (higher than 59 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.11, then +0.16, then +0.01 — a plateau around step 3, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.73 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.71 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.73, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 67 s · snippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.229 before conversion and 0.511 after — it rose by 0.282. Neighbour-to-neighbour the worst pair went 0.337 → 0.689. (The earlier render, with segment 1 left raw, scores 0.424 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.282 in the original and +0.844 after conversion — 299 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Hope Enthusiasm Optimism, -0.403 became -0.562.

Quality. Mean predicted overall quality across the segments went 2.87 → 3.18 (+0.31) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.229 → 0.511 +0.282identity cos neighbours 0.337 → 0.689d_b rescored +0.282 → +0.844d_a rescored -0.403 → -0.562d_a mined -0.403d_b mined 0.291min_cos_consec (site) 0.7073min_cos_anchor (site) 0.7340dataset snippetslang ?speaker batch275_part0_batch275_patotal 66.3schain gain +3.8 dBseam step 1.6 dBcrossfades 150/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, slightly relaxed, fairly steady, light breath
(hope enthusiasm optimism, astonishment surprise, jealousy and envy · brisk, normally alert, some disfluency, casual) being on the same team, backed by Facebook's huge wealth, seemed like a win-win. The acquisition happened in 2012. And back then, nobody had ever paid a billion dollars for a mobile app before, especially not an app that was only 18 months old and had literally no business model. Instagram wasn't making any money still. So most people thought it was absurd. Facebook were buying it for a billion dollars. Of course, in hindsight, it was a bargain. By 2019, Instagram was generating over 24 billion dollars in revenue per year, around 30% of Facebook's total revenue. Plus, it also meant that the top alternative to Facebook was now owned by Facebook.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, astonishment surprise, jealousy and envy; style: casual, monologue; good recording, quiet background; genuineness 1.8/6; vocal-burst blend 9.3/10; 35.7s.
batch275_part0_batch275_part0_chunk_947_1_780285 · in -17.1 dBFS · gain -2.9 dB · snippets-00913
(shame · brisk, energised, almost no disfluency, storytelling) But it was true. They were both leaving. In the months after the Instagram founders quit, their app was rebranded as Instagram from Facebook. The frequency of advertising on Instagram also increased, and there are now more notifications and personalized recommendations about who to follow, along with more integration with Facebook.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as shame; style: storytelling, dramatic; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 4.4/10; 16.4s.
batch275_part0_batch275_part0_chunk_947_1_780323 · in -16.4 dBFS · gain -3.6 dB · snippets-00913
(disappointment, fear · brisk, normally alert, almost no disfluency, dramatic) given how high inflation is right now, money left in normal savings accounts doesn't stand much of a chance. Whereas the average annual return across all fundraise clients in 2021 was 22.99%.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as disappointment, fear; style: dramatic, playful; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 3.0/10; 11.3s.
batch275_part0_batch275_part0_chunk_947_1_780346 · in -16.8 dBFS · gain -3.2 dB · snippets-00913
(normal-paced, normally alert, some disfluency, casual) The app was a way for people to say where they were or where they were planning to go.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 4.3/6; vocal-burst blend 5.2/10; 3.3s.
batch275_part0_batch275_part0_chunk_947_1_780452 · in -17.0 dBFS · gain -3.0 dB · snippets-00913
Longing ↓  /  Reliefidentity +0.10 emotion 136 %   c-snippets-AB2 · #14

This chain comes from the two-sided rule: it only counts if both emotions move — Longing down and Relief up — by at least 0.25 each.

The chain starts with Relief clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.25.

At the same time Longing goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.22, then -0.02, then +0.05 — not a clean run: step 2 moves back the other way by 0.02 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.72 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.77 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.72, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 17 s · snippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.711 before conversion and 0.807 after — it rose by 0.096. Neighbour-to-neighbour the worst pair went 0.768 → 0.789. (The earlier render, with segment 1 left raw, scores 0.711 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.241 in the original and +0.328 after conversion — 136 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Longing, -0.261 became -0.236.

Quality. Mean predicted overall quality across the segments went 2.88 → 2.92 (+0.04) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.711 → 0.807 +0.096identity cos neighbours 0.768 → 0.789d_b rescored +0.241 → +0.328d_a rescored -0.261 → -0.236d_a mined -0.260d_b mined 0.250min_cos_consec (site) 0.7713min_cos_anchor (site) 0.7187dataset snippetslang ?speaker batch195_part4_batch195_patotal 16.5schain gain +1.3 dBseam step 0.6 dBcrossfades 150/100/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, very good recording, no background noise, slightly relaxed, no disfluency
(longing, jealousy and envy, infatuation · normal-paced, normally alert, fairly steady, formal) Seine Augen strahlten vor Neugier und Abenteuerlust.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing, jealousy and envy, infatuation; style: formal, monologue; very good recording, no background noise; genuineness 0.5/6; vocal-burst blend 1.9/10; 3.8s.
batch195_part4_batch195_part4_chunk_2754_1_2705057 · in -23.9 dBFS · gain +3.9 dB · snippets-00496
(hope enthusiasm optimism, contentment, thankfulness gratitude · slow, subdued, fairly steady, monologue) Ich bin hier, um dir zu zeigen, dass es noch Hoffnung gibt.
full caption & clip details
An adult masculine voice; delivery is subdued, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as hope enthusiasm optimism, contentment, thankfulness gratitude; style: monologue, formal; very good recording, no background noise; genuineness 0.7/6; vocal-burst blend 7.0/10; 4.2s.
batch195_part4_batch195_part4_chunk_2754_1_2705092 · in -25.9 dBFS · gain +5.9 dB · snippets-00496
(contentment, affection, longing · slow, normally alert, steady, monologue) Lass uns gemeinsam atmen, mein Sohn und die Last deiner Sorgen teilen.
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; crisply articulate, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contentment, affection, longing; style: monologue, formal; very good recording, no background noise; genuineness 0.3/6; vocal-burst blend 2.9/10; 5.1s.
batch195_part4_batch195_part4_chunk_2754_1_2705099 · in -24.3 dBFS · gain +4.3 dB · snippets-00496
(relief, jealousy and envy, affection · measured, normally alert, fairly steady, storytelling) Aber es ist wichtig, dass du weißt, du bist nie allein.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as relief, jealousy and envy, affection; style: storytelling, narration; very good recording, no background noise; genuineness 0.6/6; vocal-burst blend 5.4/10; 3.9s.
batch195_part4_batch195_part4_chunk_2754_1_2705102 · in -26.2 dBFS · gain +6.2 dB · snippets-00496
Astonishment Surprise ↓  /  Infatuationidentity −0.06 emotion 72 %   c-snippets-AB2 · #15

This chain comes from the two-sided rule: it only counts if both emotions move — Astonishment Surprise down and Infatuation up — by at least 0.25 each.

The chain starts with Infatuation below average — 0.42, lower than 58 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.47.

At the same time Astonishment Surprise goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.03, then +0.25, then -0.05, then +0.24 — not a clean run: step 3 moves back the other way by 0.05 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 32 s · snippets

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.775 before conversion and 0.717 after — it fell by 0.059. Neighbour-to-neighbour the worst pair went 0.775 → 0.717. (The earlier render, with segment 1 left raw, scores 0.659 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.476 in the original and +0.344 after conversion — 72 % of the delta retained, which is most of it. On the other named axis, Astonishment Surprise, -0.283 became -0.447.

Quality. Mean predicted overall quality across the segments went 2.83 → 2.84 (+0.02) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.775 → 0.717 -0.059identity cos neighbours 0.775 → 0.717d_b rescored +0.476 → +0.344d_a rescored -0.283 → -0.447d_a mined -0.283d_b mined 0.472min_cos_consec (site) 0.8120min_cos_anchor (site) 0.8120dataset snippetslang ?speaker batch207_part0_batch207_patotal 30.3schain gain +1.9 dBseam step 1.1 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, good recording, no background noise, fairly steady, light breath
(astonishment surprise · brisk, normally alert, slightly relaxed, casual) And did you know that Toyota began cutting down on fossil fuel powered vehicles back in 1997 when it first rolled out the Prius?
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as astonishment surprise; style: casual, conversational; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 2.4/10; 7.1s.
batch207_part0_batch207_part0_chunk_332_1_36347 · in -20.7 dBFS · gain +0.7 dB · snippets-00555
(normal-paced, normally alert, slightly relaxed, conversational) But now, Toyota's come up with something completely different.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: conversational, casual; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 4.3/10; 3.1s.
batch207_part0_batch207_part0_chunk_332_1_36587 · in -21.3 dBFS · gain +1.3 dB · snippets-00555
(brisk, normally alert, slightly relaxed, casual) So, you may have heard about the Mirai, the hydrogen-powered Toyota vehicle that uses fuel cells to generate electricity.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 3.1/10; 6.5s.
batch207_part0_batch207_part0_chunk_332_1_36601 · in -20.6 dBFS · gain +0.6 dB · snippets-00555
(awe · brisk, energised, neutral tension, storytelling) Hydrogen is the most abundant element in the universe and has the highest specific energy density of any non-nuclear power source.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, full; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as awe; style: storytelling, dramatic; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 2.0/10; 7.6s.
batch207_part0_batch207_part0_chunk_332_1_36611 · in -21.5 dBFS · gain +1.5 dB · snippets-00555
(normal-paced, normally alert, slightly relaxed, monologue) They've added stronger connecting rods, harder valves and valve seats, and fuel injectors that use gas instead of liquid.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: monologue, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.0/10; 6.5s.
batch207_part0_batch207_part0_chunk_332_1_36637 · in -21.6 dBFS · gain +1.6 dB · snippets-00555
Bitterness ↓  /  Emotional Numbnessidentity +0.40 emotion 66 %   c-snippets-AB2 · #16

This chain comes from the two-sided rule: it only counts if both emotions move — Bitterness down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness around average — 0.57, higher than 57 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.28.

At the same time Bitterness goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.46. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.14, then -0.09 — not a clean run: step 3 moves back the other way by 0.09 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.04 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.18 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.04, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 18 s · snippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.019 before conversion and 0.380 after — it rose by 0.399. Neighbour-to-neighbour the worst pair went -0.041 → 0.380. (The earlier render, with segment 1 left raw, scores 0.237 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.280 in the original and +0.184 after conversion — 66 % of the delta retained. On the other named axis, Bitterness, -0.466 became -0.373.

Quality. Mean predicted overall quality across the segments went 2.51 → 2.64 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 -0.019 → 0.380 +0.399identity cos neighbours -0.041 → 0.380d_b rescored +0.280 → +0.184d_a rescored -0.466 → -0.373d_a mined -0.465d_b mined 0.283min_cos_consec (site) -0.1809min_cos_anchor (site) 0.0385dataset snippetslang ?speaker batch242_part0_batch242_patotal 16.8schain gain +1.8 dBseam step 2.5 dBcrossfades 150/150/100 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, normally alert, light breath
(confusion, bitterness, pain · normal-paced, neutral tension, fairly steady, casual) (low mumble) all I know is that I got hit from the side and all I'm saying is that it either was a CEO or another inmate.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as confusion, bitterness, pain; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 5.1/10; 7.2s.
batch242_part0_batch242_part0_chunk_64_1_234670 · in -22.5 dBFS · gain +2.5 dB · snippets-00742
(infatuation · measured, slightly relaxed, steady, monologue) where he's earned the nickname cheese for his ever-present smile.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, very full; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation; style: monologue, formal; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 2.5/10; 3.4s.
batch242_part0_batch242_part0_chunk_64_1_234691 · in -22.0 dBFS · gain +2.0 dB · snippets-00742
(emotional numbness · measured, slightly relaxed, fairly steady, casual) Essentially, Louisville is made up of everybody's 1%.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as emotional numbness; style: casual, conversational; average recording, no background noise; genuineness 3.1/6; vocal-burst blend 2.1/10; 3.5s.
batch242_part0_batch242_part0_chunk_64_1_234693 · in -22.1 dBFS · gain +2.1 dB · snippets-00742
(normal-paced, slightly relaxed, fairly steady, formal) And after any event, it's up to the major to follow up.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 3.4/10; 3.2s.
batch242_part0_batch242_part0_chunk_64_1_234724 · in -20.9 dBFS · gain +0.9 dB · snippets-00742
Disgust ↓  /  Contentmentidentity +0.37 emotion 283 %   c-snippets-AB2 · #17

This chain comes from the two-sided rule: it only counts if both emotions move — Disgust down and Contentment up — by at least 0.25 each.

The chain starts with Contentment clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.27.

At the same time Disgust goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.20, then -0.03, then +0.09 — not a clean run: step 2 moves back the other way by 0.03 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.05 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.04 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.05, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 47 s · snippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.104 before conversion and 0.269 after — it rose by 0.373. Neighbour-to-neighbour the worst pair went 0.075 → 0.297. (The earlier render, with segment 1 left raw, scores 0.091 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.263 in the original and +0.744 after conversion — 283 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Disgust, -0.335 became -0.450.

Quality. Mean predicted overall quality across the segments went 2.33 → 3.03 (+0.70) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 -0.104 → 0.269 +0.373identity cos neighbours 0.075 → 0.297d_b rescored +0.263 → +0.744d_a rescored -0.335 → -0.450d_a mined -0.335d_b mined 0.267min_cos_consec (site) 0.0436min_cos_anchor (site) -0.0548dataset snippetslang ?speaker batch110_part2_batch110_patotal 45.7schain gain +3.2 dBseam step 1.9 dBcrossfades 100/100/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult somewhat masculine voice · relaxed, frequent disfluency
(disgust, intoxication altered states of consciousness, interest · normal-paced, normally alert, moderately variable, casual) texture, and it's so thin, but it's not crispy like a cookie. It's it's it's a very, very, very soft.
full caption & clip details
A young adult somewhat masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is slightly cool, slightly bright, slightly rough, thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as disgust, intoxication altered states of consciousness, interest; style: casual, conversational; below-average recording, quiet background; mildly explicit content; genuineness 4.1/6; vocal-burst blend 3.0/10; 8.0s.
batch110_part2_batch110_part2_chunk_1990_1_1835787 · in -22.0 dBFS · gain +2.0 dB · snippets-00055
(teasing, relief, pleasure ecstasy · normal-paced, energised, moderately variable, casual) Wow. Hey. Yeah. Open my window. Hey, meet you a friend of ours, asked me to help you. Oh. (surprised gasp) I managed to get the key to yourself. Oh, that's lucky. I think I will climb onto the roof, so I can drop it to one of the ventilation shafts. You don't have an aerial?
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, relaxed, moderately variable; timbre is slightly cool, very dark, slightly rough, slightly thin; slurred, frequent disfluency, very wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as teasing, relief, pleasure ecstasy; style: casual, playful; below-average recording, some background noise; mildly explicit content; genuineness 5.1/6; vocal-burst blend 1.7/10; 17.6s.
batch110_part2_batch110_part2_chunk_1990_1_1836053 · in -28.7 dBFS · gain +8.7 dB · snippets-00055
(infatuation, contemplation, doubt · slow, very low-energy, steady, whispered) So there are (wistful sigh) most of the time two different styles of carpets that we offer and that you might want to consider.
full caption & clip details
An elderly feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, smooth, slightly thin; slurred, frequent disfluency, narrow pitch range, audible breath; affect is negative, submissive, vulnerable; reads as infatuation, contemplation, doubt; style: whispered, ASMR; below-average recording, no background noise; genuineness 1.7/6; vocal-burst blend 0.5/10; 9.7s.
batch110_part2_batch110_part2_chunk_1990_1_1836082 · in -32.2 dBFS · gain +12.2 dB · snippets-00055
(contentment, pleasure ecstasy, affection · slow, very low-energy, moderately variable, whispered) So, for window treatments, I brought two different styles for you to pick from, just so that we know the direction that you're willing to go to.
full caption & clip details
An elderly feminine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is slightly warm, very dark, smooth, slightly thin; slurred, frequent disfluency, narrow pitch range, audible breath; affect is mildly positive, submissive, vulnerable; reads as contentment, pleasure ecstasy, affection; style: whispered, storytelling; poor recording, quiet background; genuineness 1.6/6; vocal-burst blend 1.5/10; 10.8s.
batch110_part2_batch110_part2_chunk_1990_1_1836109 · in -31.0 dBFS · gain +11.0 dB · snippets-00055
Pride ↓  /  Affectionidentity +0.53 emotion 59 %   c-snippets-AB2 · #18

This chain comes from the two-sided rule: it only counts if both emotions move — Pride down and Affection up — by at least 0.25 each.

The chain starts with Affection clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.37.

At the same time Pride goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.67 (higher than 67 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.16 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.22 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.43 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.22, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 22 s · snippets

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.220 before conversion and 0.750 after — it rose by 0.530. Neighbour-to-neighbour the worst pair went 0.436 → 0.750. (The earlier render, with segment 1 left raw, scores 0.656 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Affection moved +0.371 in the original and +0.221 after conversion — 59 % of the delta retained. On the other named axis, Pride, -0.307 became -0.274.

Quality. Mean predicted overall quality across the segments went 2.62 → 2.91 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.220 → 0.750 +0.530identity cos neighbours 0.436 → 0.750d_b rescored +0.371 → +0.221d_a rescored -0.307 → -0.274d_a mined -0.303d_b mined 0.370min_cos_consec (site) 0.4267min_cos_anchor (site) 0.2241dataset snippetslang ?speaker batch140_part4_batch140_patotal 21.0schain gain +4.5 dBseam step 2.4 dBcrossfades 100/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · average recording, quiet background, frequent disfluency, somewhat unclear
(pride · normal-paced, subdued, slightly relaxed, monologue) a sort of one-on-one uh (low mumble) training session where we go out ourselves and work with students and teachers directly (low mumble) um and that
full caption & clip details
A young adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as pride; style: monologue, casual; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 3.8/10; 9.3s.
batch140_part4_batch140_part4_chunk_225_1_526861 · in -14.9 dBFS · gain -5.1 dB · snippets-00215
(slow, normally alert, relaxed, casual) trying to study science with all the challenges against him.
full caption & clip details
A young adult masculine voice; delivery is normally alert, slow, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 3.0/10; 5.1s.
batch140_part4_batch140_part4_chunk_225_1_526920 · in -12.7 dBFS · gain -7.3 dB · snippets-00215
(affection, infatuation, contentment · normal-paced, normally alert, relaxed, casual) (low mumble) and once we met, he was, he was a different kind of teacher, very interested.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is slightly cool, slightly bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, audible breath; affect is positive, slightly submissive, neutral openness; reads as affection, infatuation, contentment; style: casual, conversational; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 3.4/10; 6.9s.
batch140_part4_batch140_part4_chunk_225_1_526966 · in -17.2 dBFS · gain -2.8 dB · snippets-00215
Emotional Numbness ↓  /  Fearidentity +0.35 emotion 132 %   c-snippets-AB2 · #19

This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Fear up — by at least 0.25 each.

The chain starts with Fear clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.27.

At the same time Emotional Numbness goes the other way, from 0.85 (higher than 85 % of clips in this corpus) to 0.59 (higher than 59 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.12 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.06 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.04 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.06, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 18 s · snippets

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.086 before conversion and 0.433 after — it rose by 0.347. Neighbour-to-neighbour the worst pair went 0.046 → 0.404. (The earlier render, with segment 1 left raw, scores 0.344 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.269 in the original and +0.354 after conversion — 132 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, -0.258 became +0.076.

Quality. Mean predicted overall quality across the segments went 2.82 → 2.99 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.086 → 0.433 +0.347identity cos neighbours 0.046 → 0.404d_b rescored +0.269 → +0.354d_a rescored -0.258 → +0.076d_a mined -0.264d_b mined 0.269min_cos_consec (site) 0.0358min_cos_anchor (site) 0.0648dataset snippetslang ?speaker batch256_part3_batch256_patotal 17.8schain gain +2.1 dBseam step 0.5 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(normal-paced, no disfluency, clear, narration) This also required architects and engineers to come up with innovative solutions for overcoming the difficult constraints of the site.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.9/10; 6.3s.
batch256_part3_batch256_part3_chunk_777_1_1032023 · in -26.5 dBFS · gain +6.5 dB · snippets-00816
(brisk, almost no disfluency, clear, narration) Each crane was equipped with GPS, which detected the position, rotational direction, and speed of its jib in real time.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.4/10; 7.1s.
batch256_part3_batch256_part3_chunk_777_1_1032040 · in -27.3 dBFS · gain +7.3 dB · snippets-00816
(fear · normal-paced, some disfluency, average clarity, casual) formed yeah, I was virtually buried alive, inasmuch as the shopping area.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as fear; style: casual, conversational; good recording, no background noise; genuineness 2.8/6; vocal-burst blend 3.8/10; 4.8s.
batch256_part3_batch256_part3_chunk_777_1_1032232 · in -27.6 dBFS · gain +7.6 dB · snippets-00816
Contemplation ↓  /  Concentrationidentity +0.21 emotion 75 %   c-snippets-AB2 · #20

This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Concentration up — by at least 0.25 each.

The chain starts with Concentration around average — 0.56, higher than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.34.

At the same time Contemplation goes the other way, from 0.77 (higher than 77 % of clips in this corpus) to 0.51 (higher than 51 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.09, then +0.04, then +0.22 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.12 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.06 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.12, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 25 s · snippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.175 before conversion and 0.383 after — it rose by 0.209. Neighbour-to-neighbour the worst pair went 0.102 → 0.414. (The earlier render, with segment 1 left raw, scores 0.391 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.344 in the original and +0.258 after conversion — 75 % of the delta retained, which is most of it. On the other named axis, Contemplation, -0.260 became +0.037.

Quality. Mean predicted overall quality across the segments went 2.82 → 2.88 (+0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.175 → 0.383 +0.209identity cos neighbours 0.102 → 0.414d_b rescored +0.344 → +0.258d_a rescored -0.260 → +0.037d_a mined -0.263d_b mined 0.341min_cos_consec (site) 0.0636min_cos_anchor (site) 0.1237dataset snippetslang ?speaker batch69_part1_batch69_parttotal 24.0schain gain +0.1 dBseam step 2.1 dBcrossfades 100/100/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, clear
(normal-paced, fairly steady, no disfluency, formal) be a key moment in the development of quantum computer theory.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 2.5/10; 4.0s.
batch69_part1_batch69_part1_chunk_1616_1_1632428 · in -21.9 dBFS · gain +1.9 dB · snippets-01240
(measured, fairly steady, no disfluency, formal) Russian German mathematician, Yuri Minen, was the first to propose the idea of quantum computing.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.3/10; 6.3s.
batch69_part1_batch69_part1_chunk_1616_1_1632558 · in -23.2 dBFS · gain +3.2 dB · snippets-01240
(emotional numbness · normal-paced, fairly steady, little disfluency, formal) A voltage to the tip just above that silicon hydrogen bond and literally release one hydrogen atom from the surface, leaving a a dangling bond.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness; style: formal, monologue; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 1.4/10; 6.2s.
batch69_part1_batch69_part1_chunk_1616_1_1632678 · in -24.9 dBFS · gain +4.8 dB · snippets-01240
(concentration · normal-paced, steady, almost no disfluency, formal) Another significant algorithm is Grover's algorithm, which was devised to sort through information in unordered databases.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: formal, monologue; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.0/10; 8.0s.
batch69_part1_batch69_part1_chunk_1616_1_1632720 · in -23.4 dBFS · gain +3.4 dB · snippets-01240