Manifest tier. emotion_twosided, rule AB2, T=0.2, step cap 0.25. Population 1,224,717 chains (12,183 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 942,844.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words
are a real non-speech sound, printed where it happens.
The full generated caption for any clip is under “full caption & clip
details”. Its perceived-gender and background-noise clauses were re-rendered from
the numeric buckets, because the versions stored in the corpus index had those two
ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
The tags and captions describe the ORIGINAL clips — they are the
corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the
converted audio are the ones printed in each card's own paragraph and chips.
This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Fear up — by at least 0.20 each.
The chain starts with Fear around average — 0.47, lower than 53 % of clips in this corpus — and ends with it clearly present at 0.73, higher than 73 % of clips in this corpus. That is a total rise of 0.26.
At the same time Concentration goes the other way, from 0.82 (higher than 82 % of clips in this corpus) to 0.56 (higher than 56 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.04, then +0.22 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 28 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.904 before conversion and 0.819 after — it fell by 0.085. Neighbour-to-neighbour the worst pair went 0.904 → 0.859. (The earlier render, with segment 1 left raw, scores 0.693 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.261 in the original and +0.280 after conversion — 107 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.257 became -0.281.
Quality. Mean predicted overall quality across the segments went 2.97 → 3.12 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.904 → 0.819-0.085identity cos neighbours 0.904 → 0.859d_b rescored +0.261 → +0.280d_a rescored -0.257 → -0.281d_a mined -0.258d_b mined 0.261min_cos_consec (site) 0.9332min_cos_anchor (site) 0.9332dataset emolialang enspeaker EN_B00064_S09013total 27.7schain gain +1.8 dBseam step 0.7 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(normal-paced, steady, no disfluency, newsreading)This included secular works, such as the Bovo book, and religious writing specifically for women, such as the ZNH Rhein Seino-Urino and the Thangwt Tkienz
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 11.5s, EN.
EN_B00064_S09013_W000018 · in -16.4 dBFS · gain -3.6 dB · emolia-01469
(emotional numbness·measured, steady, almost no disfluency, formal)Around 5 million of those killed — 85% of the Jews who died in the Holocaust — were speakers of Yiddish
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 8.3s, EN.
EN_B00064_S09013_W000019 · in -16.5 dBFS · gain -3.5 dB · emolia-01469
(normal-paced, fairly steady, no disfluency, formal)There are well over 30,000 Yiddish speakers in the United Kingdom, and several thousand children now have Yiddish as a first language
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.3/10; 8.3s, EN.
EN_B00064_S09013_W000020 · in -17.5 dBFS · gain -2.5 dB · emolia-01469
This chain comes from the two-sided rule: it only counts if both emotions move — Bitterness down and Fear up — by at least 0.20 each.
The chain starts with Fear around average — 0.56, higher than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.39.
At the same time Bitterness goes the other way, from 0.88 (higher than 88 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.04 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.08 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.04, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 38 s · de · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.050 before conversion and 0.706 after — it rose by 0.656. Neighbour-to-neighbour the worst pair went 0.037 → 0.697. (The earlier render, with segment 1 left raw, scores 0.546 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.388 in the original and +0.299 after conversion — 77 % of the delta retained, which is most of it. On the other named axis, Bitterness, -0.316 became -0.580.
Quality. Mean predicted overall quality across the segments went 3.07 → 3.27 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.050 → 0.706+0.656identity cos neighbours 0.037 → 0.697d_b rescored +0.388 → +0.299d_a rescored -0.316 → -0.580d_a mined -0.337d_b mined 0.388min_cos_consec (site) 0.0820min_cos_anchor (site) 0.0361dataset emolialang despeaker DE_vFkmgr4cWTQtotal 37.5schain gain +2.3 dBseam step 0.6 dBcrossfades 100/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(measured, frequent disfluency, somewhat unclear, monologue)Aber die SA ist viel, viel mehr als das. In ihrer Breite ist sie eine durchaus sehr bürgerliche Organisation, die
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 0.0/10; 9.2s, DE.
DE_vFkmgr4cWTQ_W000163 · in -19.0 dBFS · gain -1.0 dB · emolia-00091
(jealousy and envy, concentration, confusion·normal-paced, frequent disfluency, average clarity, monologue)eben, sehr formalisiert auch ist. Also, es ist so unheimlich spannend zu sehen, wie viele Befehle und Erlässe die, (low mumble) ähm, SA-Führung, (low mumble) ähm, veröffentlicht, um eben die, die SA-Angehörigen zu disziplinieren und das widerspricht vollkommen dem, äh, (ahem) Bild des SA-Rabauken.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, concentration, confusion; style: monologue, formal; average recording, no background noise; genuineness 3.7/6; vocal-burst blend 0.0/10; 19.3s, DE.
DE_vFkmgr4cWTQ_W000164 · in -19.4 dBFS · gain -0.6 dB · emolia-00091
(fear· normal-paced, almost no disfluency, clear, didactic)Die SA wird dann ja 1932 von der Regierung Brüning verboten. Aber dieses Verbot wird wenig später, nach dem Regierungswechsel, wieder aufgehoben.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear; style: didactic, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 9.3s, DE.
DE_vFkmgr4cWTQ_W000165 · in -20.5 dBFS · gain +0.5 dB · emolia-00091
This chain comes from the two-sided rule: it only counts if both emotions move — Sexual Lust down and Hope Enthusiasm Optimism up — by at least 0.20 each.
The chain starts with Hope Enthusiasm Optimism clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.26.
At the same time Sexual Lust goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.35. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.04, then +0.02 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.61 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.63 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.61, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 41 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.551 before conversion and 0.641 after — it rose by 0.090. Neighbour-to-neighbour the worst pair went 0.663 → 0.714. (The earlier render, with segment 1 left raw, scores 0.453 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.265 in the original and +0.217 after conversion — 82 % of the delta retained, which is most of it. On the other named axis, Sexual Lust, -0.347 became -0.322.
Quality. Mean predicted overall quality across the segments went 2.78 → 3.07 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.551 → 0.641+0.090identity cos neighbours 0.663 → 0.714d_b rescored +0.265 → +0.217d_a rescored -0.347 → -0.322d_a mined -0.347d_b mined 0.264min_cos_consec (site) 0.6281min_cos_anchor (site) 0.6114dataset emolialang enspeaker EN_pxznqyV9TGYtotal 39.7schain gain +1.4 dBseam step 0.3 dBcrossfades 100/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · slightly rough, balanced body, energised, neutral tension, moderately variable, some disfluency, average clarity, wide pitch range
(sexual lust, contempt, teasing · brisk, light breath, casual, playful)You mean the Pirates of Penzance?
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as sexual lust, contempt, teasing; style: casual, playful; below-average recording, some background noise; genuineness 4.2/6; vocal-burst blend 3.5/10; 6.1s, EN.
EN_pxznqyV9TGY_W000001 · in -15.8 dBFS · gain -4.2 dB · emolia-02609
(jealousy and envy, teasing, malevolence malice·normal-paced, light breath, casual, conversational)Thespian! And what if this Master Thespian thing doesn't work out?
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as jealousy and envy, teasing, malevolence malice; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 3.9/6; vocal-burst blend 4.1/10; 9.3s, EN.
EN_pxznqyV9TGY_W000002 · in -15.8 dBFS · gain -4.2 dB · emolia-02609
(astonishment surprise, elation, amusement·brisk, normal breath, casual, storytelling)Set construction, electrical engineering, and I can get class credits just for being involved. Sounds like you're Broadway bound. Not Broadway bound, Redwood Road bound.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, fairly guarded; reads as astonishment surprise, elation, amusement; style: casual, storytelling; average recording, some background noise; mildly explicit content; genuineness 3.9/6; vocal-burst blend 4.5/10; 10.6s, EN.
EN_pxznqyV9TGY_W000003 · in -14.8 dBFS · gain -5.2 dB · emolia-02609
(hope enthusiasm optimism, elation · brisk, light breath, casual, playful)The Gilbert and Sullivan Festival is on the Redwood Road campus of Salt Lake Community College. What if I want to get involved? Easy, just call Salt Lake Community College at 957-4073. 957-4073? 957-4073.
full caption & clip details
A child masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, elation; style: casual, playful; average recording, quiet background; mildly explicit content; genuineness 2.0/6; vocal-burst blend 3.8/10; 14.4s, EN.
EN_pxznqyV9TGY_W000004 · in -16.0 dBFS · gain -4.0 dB · emolia-02609
This chain comes from the two-sided rule: it only counts if both emotions move — Affection down and Contemplation up — by at least 0.20 each.
The chain starts with Contemplation clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.36.
At the same time Affection goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.35. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.22, then -0.07, then +0.02, then +0.20 — not a clean run: step 2 moves back the other way by 0.07 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.37 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.40 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.37, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 133 s · sv · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.345 before conversion and 0.865 after — it rose by 0.520. Neighbour-to-neighbour the worst pair went 0.375 → 0.897. (The earlier render, with segment 1 left raw, scores 0.782 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.361 in the original and +0.479 after conversion — 133 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Affection, -0.351 became -0.075.
Quality. Mean predicted overall quality across the segments went 3.18 → 3.34 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.345 → 0.865+0.520identity cos neighbours 0.375 → 0.897d_b rescored +0.361 → +0.479d_a rescored -0.351 → -0.075d_a mined -0.353d_b mined 0.357min_cos_consec (site) 0.3973min_cos_anchor (site) 0.3743dataset podcastlang svspeaker 679300total 131.9schain gain +2.6 dBseam step 1.9 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, fairly steady, moderate pitch range, light breath
(affection, jealousy and envy, relief · normal-paced, normally alert, slightly relaxed, casual)Falsk positiv. (ahem) Som gör att han bli ariserat. Och så där kunde vi att gå upp igenare. Men poängen var med att snacka om digital dehumanisering är ju net på att autonom vapensystemer som kan rätta sig emot människor vill ju befinna sig helt extremt i än av detta digital dehumaniseringsspektra. (ahem)
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as affection, jealousy and envy, relief; style: casual, monologue; good recording, quiet background; genuineness 1.7/6; vocal-burst blend 4.6/10; 26.4s, SV.
679300_00074680 · in -41.0 dBFS · gain +21.0 dB · podcast-05163
(jealousy and envy, interest, contentment· normal-paced, normally alert, slightly relaxed, monologue)Träck med dig. Och det är en målprofil som ger match på dig, och det är ett autonomt opensystem. Så har du kun möjlighet att gå upp igen det äta på. För det har du missigt livet i sant. Så detta är det måta att peka på den reduktion av mänska i en målprofil. Och bygga den linjen för oss om att det här har vi var kunna vara eniga om att vi icka önskar i ett
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as jealousy and envy, interest, contentment; style: monologue, casual; good recording, quiet background; genuineness 2.1/6; vocal-burst blend 6.3/10; 29.5s, SV.
679300_00077312 · in -40.0 dBFS · gain +20.0 dB · podcast-05161
(contentment, relief, sexual lust· normal-paced, subdued, slightly relaxed, monologue)Det med reduktion till datapunkter är egentligen förutsättning för att netta på maskiner och processor och kanske kunstintelligens ska kunna ta avvelser. För det se och i världen som vi är. Så den förenklingen av verkligheten närhet nödvändig för att vi ska få hjälp av maskiner till att processera massa av data. Och det är ju netta i den här processen med förenklingar. (ahem) Och många lagar av förenklingar att det kan göra det för.
full caption & clip details
A young adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as contentment, relief, sexual lust; style: monologue, casual; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 9.5/10; 25.1s, SV.
679300_00080424 · in -39.4 dBFS · gain +19.4 dB · podcast-05144
(relief, interest, concentration·measured, subdued, slightly relaxed, monologue)Jag tycker att det var lagteknologi där hälgen neutral handling. Så det finns fodomar i samfunderna det. Vi har sett att maskinlärning kan vara problematisk av den lär av (ahem) det för man aldrig existerade i samfundet, och av den input den får. Det brukar de till att få något på andra sidan som reflekterar det som alla rede är sälet i samfundet av elever.
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as relief, interest, concentration; style: monologue, casual; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 9.0/10; 24.4s, SV.
679300_00082936 · in -39.2 dBFS · gain +19.2 dB · podcast-05144
(contemplation, relief, contentment· measured, normally alert, relaxed, casual)Men vår vill en regulering mot tackla utföringarna runt. Ett ein norma vapen med tanke på både med målprofiler och digital dem sering varje kampanjen säkert att man ska (low mumble) göra med detta med mänskliga målprofiler och malprofiler generellt.
full caption & clip details
An adolescent masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as contemplation, relief, contentment; style: casual, storytelling; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 5.8/10; 27.2s, SV.
679300_00085376 · in -39.3 dBFS · gain +19.3 dB · podcast-06183
This chain comes from the two-sided rule: it only counts if both emotions move — Interest down and Contempt up — by at least 0.20 each.
The chain starts with Contempt clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.25.
At the same time Interest goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.04 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.97 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.97), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 45 s · sr · eurospeech
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.941 before conversion and 0.905 after — it fell by 0.036. Neighbour-to-neighbour the worst pair went 0.941 → 0.921. (The earlier render, with segment 1 left raw, scores 0.785 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contempt moved +0.250 in the original and +0.172 after conversion — 69 % of the delta retained. On the other named axis, Interest, -0.256 became -0.587.
Quality. Mean predicted overall quality across the segments went 2.97 → 3.20 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.941 → 0.905-0.036identity cos neighbours 0.941 → 0.921d_b rescored +0.250 → +0.172d_a rescored -0.256 → -0.587d_a mined -0.256d_b mined 0.250min_cos_consec (site) 0.9623min_cos_anchor (site) 0.9688dataset eurospeechlang srspeaker serbia_serbia_2016_1263_18total 44.8schain gain +3.6 dBseam step 0.5 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · slightly cool, slightly bright, fairly smooth, thin, average recording, quiet background, brisk, neutral tension
(interest, shame, concentration · energised, some disfluency, clear, cartoonish)generalnom odredbom tačke 18) istog stava i člana Predloga zakona.Obzirom da je predlagač zakona predvideo saradnju samo za jedno krivično delo koje je predloženo u Predlogu zakona o izmenama i dopunama Krivičnog zakona i to za delo propisano odredbom (ahem) člana 138a
full caption & clip details
A child feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as interest, shame, concentration; style: cartoonish, ranting; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 4.8/10; 17.2s, SR.
serbia_serbia_2016_1263_18112016_5209216_5226368 · in -27.6 dBFS · gain +7.6 dB · eurospeech-02719
(shame, distress, bitterness· energised, almost no disfluency, clear, ranting)Predloga zakona o izmenama i dopunama Krivičnog zakona, nejasno je zašto je predlagač Zakona o sprečavanju nasilja u porodici nije posebno predvideo saradnji i za krivična dela iz čl. 121a
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as shame, distress, bitterness; style: ranting, dramatic; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 3.5/10; 12.7s, SR.
serbia_serbia_2016_1263_18112016_5226368_5239056 · in -28.0 dBFS · gain +8.0 dB · eurospeech-02719
(contempt, shame, bitterness ·normally alert, some disfluency, average clarity, cartoonish)(low mumble) (ahem) 187a.Obzirom da predmetna krivična dela imaju kao objekt zaštite upravo ona dobra koja mogu biti ugrožena radnjama porodičnog nasilja koja predlagač upravo predlaže da neutrališe odredbom predloženog (ahem) Predloga zakona.Takođe
full caption & clip details
A child feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as contempt, shame, bitterness; style: cartoonish, ranting; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 6.6/10; 15.3s, SR.
serbia_serbia_2016_1263_18112016_5239056_5254400 · in -26.9 dBFS · gain +6.9 dB · eurospeech-02719
This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Infatuation up — by at least 0.20 each.
The chain starts with Infatuation around average — 0.51, right about the corpus median — and ends with it strongly present at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.26.
At the same time Concentration goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.05 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 26 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.851 before conversion and 0.744 after — it fell by 0.106. Neighbour-to-neighbour the worst pair went 0.853 → 0.744. (The earlier render, with segment 1 left raw, scores 0.666 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.262 in the original and +0.179 after conversion — 68 % of the delta retained. On the other named axis, Concentration, -0.281 became -0.315.
Quality. Mean predicted overall quality across the segments went 2.99 → 3.04 (+0.05) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.851 → 0.744-0.106identity cos neighbours 0.853 → 0.744d_b rescored +0.262 → +0.179d_a rescored -0.281 → -0.315d_a mined -0.281d_b mined 0.262min_cos_consec (site) 0.9411min_cos_anchor (site) 0.9117dataset emolialang enspeaker EN_9oFM0mE5laYtotal 25.3schain gain +0.4 dBseam step 0.5 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, steady
(concentration, hope enthusiasm optimism · normal-paced, fairly narrow pitch, formal, authoritative)Paleontology, ethology, anthropology, and biogeography as well as historical approaches to developmental biology
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, hope enthusiasm optimism; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 9.6s, EN.
EN_9oFM0mE5laY_W000004 · in -15.6 dBFS · gain -4.4 dB · emolia-02454
(measured, moderate pitch range, newsreading, formal)Genomics, physiology, ecology and many other areas of the biological sciences.The comparative approach also has numerous applications in human
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.8s, EN.
EN_9oFM0mE5laY_W000005 · in -15.1 dBFS · gain -4.9 dB · emolia-02454
(normal-paced, fairly narrow pitch, formal, authoritative)Genetics, Biomedicine, and Conservation Biology
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.5/10; 4.3s, EN.
EN_9oFM0mE5laY_W000006 · in -15.2 dBFS · gain -4.8 dB · emolia-02454
Intoxication Altered States of Consciousness ↓ / Doubt ↑identity +0.08emotion 262 % emotion_twosided__AB2__T0.20__C0.25__INTERNAL · #7
This chain comes from the two-sided rule: it only counts if both emotions move — Intoxication Altered States of Consciousness down and Doubt up — by at least 0.20 each.
The chain starts with Doubt clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.26.
At the same time Intoxication Altered States of Consciousness goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.08, then -0.01, then -0.05 — not a clean run: step 3 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.70 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.78 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.70, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 47 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.700 before conversion and 0.776 after — it rose by 0.076. Neighbour-to-neighbour the worst pair went 0.792 → 0.796. (The earlier render, with segment 1 left raw, scores 0.681 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.276 in the original and +0.723 after conversion — 262 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Intoxication Altered States of Consciousness, -0.270 became -0.324.
Quality. Mean predicted overall quality across the segments went 2.77 → 3.09 (+0.32) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.700 → 0.776+0.076identity cos neighbours 0.792 → 0.796d_b rescored +0.276 → +0.723d_a rescored -0.270 → -0.324d_a mined -0.270d_b mined 0.261min_cos_consec (site) 0.7822min_cos_anchor (site) 0.6996dataset emolialang enspeaker EN_SgZtx0KfKF4total 45.7schain gain +3.4 dBseam step 1.1 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice
(intoxication altered states of consciousness, fatigue exhaustion, teasing · normal-paced, energised, neutral tension, casual)Out of Explorer! The Fox! Happy Fox Day! I think National Fox Day was yesterday.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, slightly dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, slightly dominant, slightly guarded; reads as intoxication altered states of consciousness, fatigue exhaustion, teasing; style: casual; below-average recording, some background noise; genuineness 3.8/6; vocal-burst blend 0.9/10; 9.3s, EN.
EN_SgZtx0KfKF4_W000103 · in -18.4 dBFS · gain -1.6 dB · emolia-01896
(sexual lust, pain, confusion· normal-paced, normally alert, slightly relaxed, casual)So when I hit the Z button, and it takes me like back and forth between exits, I know I have fully explored. I don't remember. Do you have to eat in this game?
full caption & clip details
A young adult somewhat masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as sexual lust, pain, confusion; style: casual, monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 2.9/10; 9.0s, EN.
EN_SgZtx0KfKF4_W000104 · in -15.3 dBFS · gain -4.7 dB · emolia-01896
(doubt, confusion, pain ·measured, very low-energy, relaxed, casual)And I'm not sure I super understand. (exhausted groan) Uh, let's see. Why? So, I can change my staff? My staff's nature? I would like to call on...
full caption & clip details
A child somewhat masculine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is slightly cool, dark, slightly rough, slightly thin; slurred, frequent disfluency, narrow pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as doubt, confusion, pain; style: casual, conversational; below-average recording, quiet background; genuineness 3.6/6; vocal-burst blend 2.6/10; 12.5s, EN.
EN_SgZtx0KfKF4_W000107 · in -16.8 dBFS · gain -3.2 dB · emolia-01896
(confusion, doubt, embarrassment· measured, normally alert, relaxed, casual)Fire. So I think it's a fire staff now? I think that's how that works? Okay. (wistful sigh) Uh, ooh, total.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, relaxed, moderately variable; timbre is neutral-toned, dark, fairly smooth, slightly thin; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as confusion, doubt, embarrassment; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 1.5/10; 9.5s, EN.
EN_SgZtx0KfKF4_W000108 · in -20.9 dBFS · gain +0.9 dB · emolia-01896
(doubt · measured, normally alert, slightly relaxed, conversational)(low mumble) Move that to actually wield that. So, I now have a new totem power here.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as doubt; style: conversational, casual; good recording, quiet background; genuineness 3.2/6; vocal-burst blend 1.8/10; 6.2s, EN.
EN_SgZtx0KfKF4_W000109 · in -17.7 dBFS · gain -2.3 dB · emolia-01896
This chain comes from the two-sided rule: it only counts if both emotions move — Contentment down and Anger up — by at least 0.20 each.
The chain starts with Anger clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.29.
At the same time Contentment goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.58 (higher than 58 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.05 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 46 s · it · eurospeech
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.910 before conversion and 0.891 after — it fell by 0.018. Neighbour-to-neighbour the worst pair went 0.910 → 0.891. (The earlier render, with segment 1 left raw, scores 0.713 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Anger moved +0.291 in the original and +0.237 after conversion — 81 % of the delta retained, which is most of it. On the other named axis, Contentment, -0.396 became -0.350.
Quality. Mean predicted overall quality across the segments went 3.10 → 3.30 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.910 → 0.891-0.018identity cos neighbours 0.910 → 0.891d_b rescored +0.291 → +0.237d_a rescored -0.396 → -0.350d_a mined -0.396d_b mined 0.291min_cos_consec (site) 0.9255min_cos_anchor (site) 0.9255dataset eurospeechlang itspeaker italy_16_345total 45.8schain gain +3.1 dBseam step 0.2 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an elderly masculine voice · slightly cool, balanced body, average recording, quiet background, normally alert, neutral tension, moderately variable
(contentment, elation, pride · measured, frequent disfluency, slurred, storytelling)che l'arbitro può decidere anche senza rispettare le norme inderogabili del diritto del lavoro. Il richiamo presente nel testo ai principi generali del diritto è comunque utile ma
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is slightly cool, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as contentment, elation, pride; style: storytelling, playful; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 5.7/10; 16.1s, IT.
italy_16_345_9076256_9092384 · in -39.9 dBFS · gain +19.9 dB · eurospeech-01683
(sourness, bitterness, anger·fast, some disfluency, average clarity, cartoonish)senza la protezione storica del diritto di lavoro, bene, si affrontino - come si è già fatto, del resto - le singole norme e si dica, ad esempio: «Questa norma è troppo rigida, la rendiamo più morbida, la deregoliamo». È successo
full caption & clip details
A child masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sourness, bitterness, anger; style: cartoonish, dramatic; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 6.1/10; 12.5s, IT.
italy_16_345_9121759_9134224 · in -40.1 dBFS · gain +20.1 dB · eurospeech-01683
(anger, triumph, shame·brisk, some disfluency, average clarity, casual)delle norme del diritto del lavoro. Questo è assolutamente inaccettabile. L'altro argomento che si avanza è: il lavoratore, in fondo, non è così un poveretto. Beh, no, non è un poveretto, però tutto il diritto del lavoro, almeno nei suoi aspetti fondamentali, è nato sul presupposto
full caption & clip details
A child masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as anger, triumph, shame; style: casual, cartoonish; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 8.8/10; 17.6s, IT.
italy_16_345_9153360_9170928 · in -40.6 dBFS · gain +20.6 dB · eurospeech-01683
This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Emotional Numbness up — by at least 0.20 each.
The chain starts with Emotional Numbness clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.27.
At the same time Concentration goes the other way, from 0.83 (higher than 83 % of clips in this corpus) to 0.31 (lower than 69 % of clips in this corpus), a change of -0.52. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.14, then -0.11, then +0.23 — not a clean run: step 2 moves back the other way by 0.11 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 24 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.746 before conversion and 0.755 after — it rose by 0.009. Neighbour-to-neighbour the worst pair went 0.835 → 0.735. (The earlier render, with segment 1 left raw, scores 0.459 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.269 in the original and +0.276 after conversion — 103 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.518 became -0.595.
Quality. Mean predicted overall quality across the segments went 2.85 → 2.95 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.746 → 0.755+0.009identity cos neighbours 0.835 → 0.735d_b rescored +0.269 → +0.276d_a rescored -0.518 → -0.595d_a mined -0.517d_b mined 0.269min_cos_consec (site) 0.9220min_cos_anchor (site) 0.9152dataset emolialang enspeaker EN_B00066_S00727total 23.1schain gain +1.6 dBseam step 0.8 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fairly steady, formal, newsreading)To redress the matter, Congress passed the Reconstruction Acts of 1867, dissolving rebel state governments and dividing the South into military districts
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 9.0s, EN.
EN_B00066_S00727_W000043 · in -15.8 dBFS · gain -4.2 dB · emolia-01515
(steady, formal, authoritative)On October 8, 1873, Baxter filed a plea of non-jurisdiction, but he believed that the court might decide against him
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 7.2s, EN.
EN_B00066_S00727_W000044 · in -15.8 dBFS · gain -4.2 dB · emolia-01515
(fairly steady, formal, authoritative)Brooks, on the other hand, had the support of the district court
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.9/10; 3.3s, EN.
EN_B00066_S00727_W000045 · in -14.6 dBFS · gain -5.4 dB · emolia-01515
(emotional numbness· fairly steady, formal, monologue)However, the minstrels would soon turn on Baxter for not following the party line
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.6/10; 4.1s, EN.
EN_B00066_S00727_W000046 · in -14.8 dBFS · gain -5.2 dB · emolia-01515
This chain comes from the two-sided rule: it only counts if both emotions move — Pride down and Contemplation up — by at least 0.20 each.
The chain starts with Contemplation clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.26.
At the same time Pride goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.10 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 32 s · de · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.884 before conversion and 0.747 after — it fell by 0.137. Neighbour-to-neighbour the worst pair went 0.884 → 0.732. (The earlier render, with segment 1 left raw, scores 0.734 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.254 in the original and +0.205 after conversion — 81 % of the delta retained, which is most of it. On the other named axis, Pride, -0.381 became -0.104.
Quality. Mean predicted overall quality across the segments went 2.95 → 3.22 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.884 → 0.747-0.137identity cos neighbours 0.884 → 0.732d_b rescored +0.254 → +0.205d_a rescored -0.381 → -0.104d_a mined -0.381d_b mined 0.255min_cos_consec (site) 0.9378min_cos_anchor (site) 0.9332dataset emolialang despeaker DE_WiD5Do8Mq-4total 30.9schain gain +4.1 dBseam step 0.8 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(pride · some disfluency, average clarity, didactic, casual)Man sollte sich realistische Ziele setzen, dann kann man auch erfolgreich sein. Wenn man sich jetzt das Ziel setzt und sagt,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride; style: didactic, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 0.0/10; 6.8s, DE.
DE_WiD5Do8Mq-4_W000065 · in -23.1 dBFS · gain +3.1 dB · emolia-00215
(doubt, interest· some disfluency, average clarity, monologue, casual)100.000 Exemplare im ersten Jahr verkaufen, dann kann man sich schon einmal sicher sein, das wird nicht, man wird für sich selbst nicht erfolgreich sein, das wird man nicht zustande bringen. Wenn man aber sagt, man setzt sich eben Ziele, ich möchte jetzt einmal Feedback von diesem Preis bekommen und natürlich hat mir das geholfen, (low mumble)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as doubt, interest; style: monologue, casual; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 2.0/10; 16.8s, DE.
DE_WiD5Do8Mq-4_W000066 · in -22.3 dBFS · gain +2.3 dB · emolia-00215
(contemplation, affection·frequent disfluency, somewhat unclear, didactic, monologue)Wenn man sich solche Ziele setzt, dann ist es auch schon am Anfang möglich, als Selfpublisher erfolgreich zu sein und
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, affection; style: didactic, monologue; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 0.0/10; 7.8s, DE.
DE_WiD5Do8Mq-4_W000067 · in -22.1 dBFS · gain +2.1 dB · emolia-00215
This chain comes from the two-sided rule: it only counts if both emotions move — Pleasure Ecstasy down and Fatigue Exhaustion up — by at least 0.20 each.
The chain starts with Fatigue Exhaustion clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.33.
At the same time Pleasure Ecstasy goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.67 (higher than 67 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.08, then +0.25 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.64 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.64 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.64, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 19 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.592 before conversion and 0.488 after — it fell by 0.103. Neighbour-to-neighbour the worst pair went 0.533 → 0.633. (The earlier render, with segment 1 left raw, scores 0.504 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.328 in the original and +0.428 after conversion — 130 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Pleasure Ecstasy, -0.329 became -0.441.
Quality. Mean predicted overall quality across the segments went 2.35 → 2.72 (+0.37) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.592 → 0.488-0.103identity cos neighbours 0.533 → 0.633d_b rescored +0.328 → +0.428d_a rescored -0.329 → -0.441d_a mined -0.329d_b mined 0.328min_cos_consec (site) 0.6425min_cos_anchor (site) 0.6367dataset emolialang enspeaker EN_nOOGJD7Fpcytotal 18.0schain gain +3.9 dBseam step 2.0 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · fairly steady
(pleasure ecstasy, intoxication altered states of consciousness, elation · measured, very low-energy, relaxed, casual)I like this one too though. I like, I like both of them though. I like both. I like both versions of him. Even the version when he-
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is slightly cool, dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, neutral openness; reads as pleasure ecstasy, intoxication altered states of consciousness, elation; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 6.0/6; vocal-burst blend 3.2/10; 9.2s, EN.
EN_nOOGJD7Fpcy_W000095 · in -18.2 dBFS · gain -1.8 dB · emolia-02524
(confusion, intoxication altered states of consciousness, embarrassment·normal-paced, normally alert, slightly relaxed, casual)Okay, I got lucky this time. I got lucky as well.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as confusion, intoxication altered states of consciousness, embarrassment; style: casual, conversational; average recording, quiet background; genuineness 5.8/6; vocal-burst blend 4.6/10; 3.5s, EN.
EN_nOOGJD7Fpcy_W000096 · in -21.7 dBFS · gain +1.7 dB · emolia-02524
(fatigue exhaustion, intoxication altered states of consciousness, pain·measured, normally alert, slightly relaxed, casual)I'll probably make post more videos, I'll probably make videos about how to make, how to make
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, thin; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, intoxication altered states of consciousness, pain; style: casual, monologue; below-average recording, quiet background; genuineness 4.1/6; vocal-burst blend 2.4/10; 5.7s, EN.
EN_nOOGJD7Fpcy_W000100 · in -19.4 dBFS · gain -0.6 dB · emolia-02524
This chain comes from the two-sided rule: it only counts if both emotions move — Contempt down and Jealousy and Envy up — by at least 0.20 each.
The chain starts with Jealousy and Envy around average — 0.57, higher than 57 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.39.
At the same time Contempt goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.15, then +0.11, then +0.13 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.31 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.37 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.31, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 67 s · no · eurospeech
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.346 before conversion and 0.717 after — it rose by 0.371. Neighbour-to-neighbour the worst pair went 0.377 → 0.802. (The earlier render, with segment 1 left raw, scores 0.562 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Jealousy and Envy moved +0.395 in the original and +0.760 after conversion — 192 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contempt, -0.376 became -0.349.
Quality. Mean predicted overall quality across the segments went 3.15 → 3.46 (+0.31) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.346 → 0.717+0.371identity cos neighbours 0.377 → 0.802d_b rescored +0.395 → +0.760d_a rescored -0.376 → -0.349d_a mined -0.376d_b mined 0.395min_cos_consec (site) 0.3736min_cos_anchor (site) 0.3089dataset eurospeechlang nospeaker norway_10092-1total 66.3schain gain +2.6 dBseam step 2.4 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-toned, balanced body, quiet background, measured, slightly relaxed, wide pitch range
(contempt, malevolence malice, concentration · energised, moderately variable, frequent disfluency, cartoonish)Hvordan vil hun koordinere regjeringens innsats i kampen mot vold og overgrep mot barn og sørge for at tiltak ikke utsettes eller skyfles mellom departementene,
full caption & clip details
A middle-aged masculine voice; delivery is energised, measured, slightly relaxed, moderately variable; timbre is neutral-toned, slightly dark, rough, balanced body; average clarity, frequent disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, fairly guarded; reads as contempt, malevolence malice, concentration; style: cartoonish, storytelling; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 0.7/10; 13.8s, NO.
norway_10092-1_18394559_18408368 · in -24.1 dBFS · gain +4.1 dB · eurospeech-02172
(anger, disgust·normally alert, fairly steady, almost no disfluency, authoritative)men i stedet gis høyeste prioritet og gjennomføres effektivt framover for barnas skyld?
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, almost no disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as anger, disgust; style: authoritative, monologue; good recording, quiet background; genuineness 0.5/6; vocal-burst blend 0.0/10; 17.2s, NO.
norway_10092-1_18408368_18425607 · in -26.6 dBFS · gain +6.6 dB · eurospeech-02172
(emotional numbness, triumph, fatigue exhaustion·subdued, fairly steady, frequent disfluency, didactic)Over de siste årene har vi sett en økning av vold og overgrep mot barn i Norge, og politiet melder om store mørketall.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is neutral, neutral stance, fairly guarded; reads as emotional numbness, triumph, fatigue exhaustion; style: didactic, monologue; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.3/10; 18.2s, NO.
norway_10092-1_18425607_18443792 · in -26.9 dBFS · gain +6.9 dB · eurospeech-02172
(jealousy and envy, pride, anger·normally alert, moderately variable, frequent disfluency, cartoonish)Også i barnehage og skole har vi i økende grad sett avdekking av overgrep gjort mot barn. Bladet Utdanning kartla i sitt januarnummer (low mumble) i år saker i barnehage og skole i foregående år.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as jealousy and envy, pride, anger; style: cartoonish, didactic; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.1/10; 17.6s, NO.
norway_10092-1_18443792_18461440 · in -24.6 dBFS · gain +4.6 dB · eurospeech-02172
This chain comes from the two-sided rule: it only counts if both emotions move — Shame down and Triumph up — by at least 0.20 each.
The chain starts with Triumph clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.29.
At the same time Shame goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.14, then +0.01, then +0.14 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.10 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.08 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.10, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 55 s · hr · eurospeech
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.075 before conversion and 0.728 after — it rose by 0.653. Neighbour-to-neighbour the worst pair went 0.078 → 0.760. (The earlier render, with segment 1 left raw, scores 0.571 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Triumph moved +0.291 in the original and +0.755 after conversion — 259 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Shame, -0.282 became -0.246.
Quality. Mean predicted overall quality across the segments went 3.07 → 3.37 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.075 → 0.728+0.653identity cos neighbours 0.078 → 0.760d_b rescored +0.291 → +0.755d_a rescored -0.282 → -0.246d_a mined -0.282d_b mined 0.291min_cos_consec (site) 0.0816min_cos_anchor (site) 0.0954dataset eurospeechlang hrspeaker croatia_20101103093129-966total 53.5schain gain +1.6 dBseam step 1.7 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · balanced body, quiet background, fairly steady
(shame, pride · normal-paced, normally alert, slightly relaxed, monologue)pri čemu nije bio kriv samo onaj koji mi je trebao izdati tu lokacijsku dozvolu, nego uvjeti da bi ja mogao podnijeti zahtjev, nesređeno stanje katastra iz 1927.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as shame, pride; style: monologue, dramatic; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 5.0/10; 11.8s, HR.
croatia_20101103093129-9663_6766784_6778544 · in -18.1 dBFS · gain -1.9 dB · eurospeech-01319
(disappointment, bitterness, sourness· normal-paced, normally alert, slightly relaxed, monologue)1927. godine na mjestu gdje je naša kuća još uvijek je ucrtana cesta. Za to zakon nije kriv, nego je kriv netko tko to nije naravno sproveo. I sad se tek ta cesta ucrtala koja postoji već eto 80 godina.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, bitterness, sourness; style: monologue, authoritative; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 1.0/10; 15.0s, HR.
croatia_20101103093129-9663_6778544_6793568 · in -21.5 dBFS · gain +1.5 dB · eurospeech-01319
(fatigue exhaustion·measured, very low-energy, relaxed, monologue)Hvala. Na redu je gospodin zastupnik Anđelko Mihalić. Izvolite. **Mihalić, Anđelko (HDZ)** Hvala lijepo gospodine predsjedniče, gospodine državni tajniče, kolegice i kolege zastupnici.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion; style: monologue; below-average recording, quiet background; genuineness 3.6/6; vocal-burst blend 0.2/10; 14.3s, HR.
croatia_20101103093129-9663_6793568_6807824 · in -21.0 dBFS · gain +1.0 dB · eurospeech-01319
(triumph, thankfulness gratitude· measured, normally alert, slightly relaxed, monologue)Evo Zakon o postupanju i uvjetima gradnje radi poticanja i ulaganja stupio je na snagu 25. lipnja 2009. godine, i danas smo čuli iz (ahem) diskusija mojih kolega
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as triumph, thankfulness gratitude; style: monologue, authoritative; below-average recording, quiet background; genuineness 2.9/6; vocal-burst blend 2.1/10; 13.1s, HR.
croatia_20101103093129-9663_6807824_6820880 · in -15.9 dBFS · gain -4.1 dB · eurospeech-01319
This chain comes from the two-sided rule: it only counts if both emotions move — Sexual Lust down and Doubt up — by at least 0.20 each.
The chain starts with Doubt clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.25.
At the same time Sexual Lust goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.10, then +0.12, then +0.04 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.57 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.69 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.57, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 37 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.445 before conversion and 0.321 after — it fell by 0.124. Neighbour-to-neighbour the worst pair went 0.445 → 0.321. (The earlier render, with segment 1 left raw, scores 0.367 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.262 in the original and +0.195 after conversion — 74 % of the delta retained, which is most of it. On the other named axis, Sexual Lust, -0.335 became -0.722.
Quality. Mean predicted overall quality across the segments went 2.64 → 2.69 (+0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.445 → 0.321-0.124identity cos neighbours 0.445 → 0.321d_b rescored +0.262 → +0.195d_a rescored -0.335 → -0.722d_a mined -0.325d_b mined 0.255min_cos_consec (site) 0.6874min_cos_anchor (site) 0.5750dataset podcastlang enspeaker 682194total 36.4schain gain +1.6 dBseam step 3.6 dBcrossfades 150/150/100 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a child masculine voice · slow, very low-energy, frequent disfluency
(sexual lust, affection, relief · slightly relaxed, volatile, crisply articulate, casual)meet what you need.
full caption & clip details
A child masculine voice; delivery is very low-energy, slow, slightly relaxed, volatile; timbre is slightly cool, dark, fairly smooth, thin; crisply articulate, frequent disfluency, narrow pitch range, audible breath; affect is mildly positive, slightly dominant, slightly guarded; reads as sexual lust, affection, relief; style: casual, conversational; poor recording, no background noise; genuineness 3.2/6; vocal-burst blend 5.3/10; 3.2s, EN.
682194_00137520 · in -21.0 dBFS · gain +1.0 dB · podcast-01525
(contemplation, infatuation, pain·relaxed, moderately variable, somewhat unclear, ASMR)And and in talking about it, you Then can share. You know, I I want to also share with you, I'm a little worried (ahem) um that that you've lost some weight. And I don't maybe it's worry, or maybe it's on purpose. Can you tell me about it? Be
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, submissive, slightly vulnerable; reads as contemplation, infatuation, pain; style: ASMR, whispered; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 1.4/10; 21.0s, EN.
682194_00137836 · in -21.2 dBFS · gain +1.2 dB · podcast-03179
(doubt, relief, affection· relaxed, steady, average clarity, whispered)curious. See if she can just say, Well, I've been trying, or or whatever she might tell. Yeah. Okay. Yeah. That makes sense.
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; average clarity, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, slightly submissive, slightly vulnerable; reads as doubt, relief, affection; style: whispered, ASMR; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 0.1/10; 9.1s, EN.
682194_00139936 · in -23.1 dBFS · gain +3.1 dB · podcast-01521
(doubt, fear, pain· relaxed, moderately variable, average clarity, conversational)Are you what are you afraid of with that?
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is neutral-toned, dark, slightly rough, thin; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly vulnerable; reads as doubt, fear, pain; style: conversational, casual; good recording, no background noise; genuineness 3.0/6; vocal-burst blend 4.3/10; 3.6s, EN.
682194_00140872 · in -21.0 dBFS · gain +1.0 dB · podcast-01526
This chain comes from the two-sided rule: it only counts if both emotions move — Fatigue Exhaustion down and Concentration up — by at least 0.20 each.
The chain starts with Concentration around average — 0.47, lower than 53 % of clips in this corpus — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.37.
At the same time Fatigue Exhaustion goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.59 (higher than 59 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.13 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 30 s · zh · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.784 before conversion and 0.864 after — it rose by 0.080. Neighbour-to-neighbour the worst pair went 0.858 → 0.890. (The earlier render, with segment 1 left raw, scores 0.718 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.364 in the original and +0.251 after conversion — 69 % of the delta retained. On the other named axis, Fatigue Exhaustion, -0.292 became -0.612.
Quality. Mean predicted overall quality across the segments went 2.94 → 3.30 (+0.36) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.784 → 0.864+0.080identity cos neighbours 0.858 → 0.890d_b rescored +0.364 → +0.251d_a rescored -0.292 → -0.612d_a mined -0.292d_b mined 0.366min_cos_consec (site) 0.8572min_cos_anchor (site) 0.8416dataset emolialang zhspeaker ZH_B00060_S05038total 29.1schain gain +1.5 dBseam step 1.0 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, normal-paced, normally alert, slightly relaxed
This chain comes from the two-sided rule: it only counts if both emotions move — Interest down and Concentration up — by at least 0.20 each.
The chain starts with Concentration clearly present — 0.58, higher than 58 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.35.
At the same time Interest goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.63 (higher than 63 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.40 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.40 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.40, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 36 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.413 before conversion and 0.658 after — it rose by 0.245. Neighbour-to-neighbour the worst pair went 0.413 → 0.658. (The earlier render, with segment 1 left raw, scores 0.663 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.351 in the original and +0.442 after conversion — 126 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Interest, -0.332 became -0.205.
Quality. Mean predicted overall quality across the segments went 2.61 → 3.12 (+0.50) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.413 → 0.658+0.245identity cos neighbours 0.413 → 0.658d_b rescored +0.351 → +0.442d_a rescored -0.332 → -0.205d_a mined -0.331d_b mined 0.350min_cos_consec (site) 0.4042min_cos_anchor (site) 0.4042dataset emolialang enspeaker EN_Yvr__f7bY6Ytotal 35.4schain gain +2.8 dBseam step 0.5 dBcrossfades 150/100 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, balanced body, quiet background, fairly steady
(interest · normal-paced, normally alert, slightly relaxed, casual)What can you tell us about the water quality that you're using here? Because the water here in the Philippines is not perfect for coffee brewing (low mumble) as it is.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest; style: casual, conversational; good recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.6/10; 8.6s, EN.
EN_Yvr__f7bY6Y_W000079 · in -16.6 dBFS · gain -3.4 dB · emolia-02096
(intoxication altered states of consciousness·measured, subdued, relaxed, casual)So (low mumble) uhm, the water here in the Philippines, it differs from cities. (low mumble) In Cebu alone, where our raw water is probably ranging around 400 (low mumble) to 500 ppm. (low mumble) Uhm, that's...
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, slightly guarded; reads as intoxication altered states of consciousness; style: casual, monologue; below-average recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.0/10; 15.0s, EN.
EN_Yvr__f7bY6Y_W000080 · in -18.0 dBFS · gain -2.0 dB · emolia-02096
(concentration· measured, subdued, relaxed, casual)The range, so that's why we needed to have a lot of, (low mumble) um, process and we also kinda, (low mumble) uh, if you're aiming to serve specialty coffee, you know, you need to
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration; style: casual, monologue; below-average recording, quiet background; genuineness 3.2/6; vocal-burst blend 1.7/10; 12.1s, EN.
EN_Yvr__f7bY6Y_W000081 · in -19.9 dBFS · gain -0.1 dB · emolia-02096
This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Emotional Numbness up — by at least 0.20 each.
The chain starts with Emotional Numbness clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.26.
At the same time Concentration goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.17, then -0.14, then +0.23 — not a clean run: step 2 moves back the other way by 0.14 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 32 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.833 before conversion and 0.762 after — it fell by 0.071. Neighbour-to-neighbour the worst pair went 0.833 → 0.762. (The earlier render, with segment 1 left raw, scores 0.720 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.260 in the original and +0.249 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.329 became -0.263.
Quality. Mean predicted overall quality across the segments went 2.96 → 3.13 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.833 → 0.762-0.071identity cos neighbours 0.833 → 0.762d_b rescored +0.260 → +0.249d_a rescored -0.329 → -0.263d_a mined -0.329d_b mined 0.260min_cos_consec (site) 0.8956min_cos_anchor (site) 0.9387dataset emolialang enspeaker EN_1gkX-KziJNMtotal 31.5schain gain +1.3 dBseam step 0.8 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(almost no disfluency, newsreading, formal)Northern Line – The signalling system on the Northern Line has also been replaced to increase capacity on the line by 20%, as the line now runs 24 trains per hour at peak times, compared to 20 previously
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 12.7s, EN.
EN_1gkX-KziJNM_W000245 · in -17.1 dBFS · gain -3.0 dB · emolia-01161
(fear·no disfluency, formal, newsreading)Capacity can be increased further if the operation of the charring cross and bank branches are separated
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear; style: formal, newsreading; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.8/10; 5.5s, EN.
EN_1gkX-KziJNM_W000246 · in -17.6 dBFS · gain -2.4 dB · emolia-01161
(no disfluency, formal, authoritative)To enable this up to 50 additional trains will be built in addition to the current 106
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.2/10; 6.7s, EN.
EN_1gkX-KziJNM_W000247 · in -16.6 dBFS · gain -3.4 dB · emolia-01161
(no disfluency, formal, newsreading)The five trains will be required for the proposed northern line extension and 45 to increase frequencies on the rest of the line
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.1/10; 7.1s, EN.
EN_1gkX-KziJNM_W000248 · in -17.2 dBFS · gain -2.8 dB · emolia-01161
This chain comes from the two-sided rule: it only counts if both emotions move — Pride down and Emotional Numbness up — by at least 0.20 each.
The chain starts with Emotional Numbness clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.26.
At the same time Pride goes the other way, from 0.85 (higher than 85 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.05, then +0.12, then +0.03, then +0.06 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 70 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.874 before conversion and 0.778 after — it fell by 0.096. Neighbour-to-neighbour the worst pair went 0.914 → 0.822. (The earlier render, with segment 1 left raw, scores 0.618 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.258 in the original and +0.313 after conversion — 121 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Pride, -0.607 became -0.334.
Quality. Mean predicted overall quality across the segments went 3.05 → 3.21 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.874 → 0.778-0.096identity cos neighbours 0.914 → 0.822d_b rescored +0.258 → +0.313d_a rescored -0.607 → -0.334d_a mined -0.321d_b mined 0.258min_cos_consec (site) 0.9594min_cos_anchor (site) 0.9590dataset emolialang enspeaker EN_xuKHyEYNWcMtotal 68.2schain gain +1.9 dBseam step 0.5 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normal-paced, normally alert, slightly relaxed
(fairly steady, almost no disfluency, minimal breath, newsreading)President Saleh, who controlled the state, Major General Ali Mohsen al-Amar, who controlled the largest share of the Republic of Yemen armed forces, and Abdullah ibn Husayn al-Amar, figurehead of the Islamist al-Islah party and Saudi Arabia's chosen broker of transnational patronage payments to various political players, including tribal sheikhs.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 19.2s, EN.
EN_xuKHyEYNWcM_W000020 · in -15.5 dBFS · gain -4.5 dB · emolia-02260
(concentration, contempt, disgust·steady, no disfluency, light breath, newsreading)The Saudi payments have been intended to facilitate the tribe's autonomy from the Yemeni government and to give the Saudi government a mechanism with which to weigh in on Yemen's political decision-making.It is a member of the United Nations, Arab League, Organization of the Islamic Cooperation
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, contempt, disgust; style: newsreading, authoritative; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 14.6s, EN.
EN_xuKHyEYNWcM_W000021 · in -14.4 dBFS · gain -5.6 dB · emolia-02260
(steady, no disfluency, light breath, authoritative)G77, Non-Aligned Movement, Arab Satellite Communications Organization, Arab Monetary Fund and the World Federation of Trade Unions
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 8.3s, EN.
EN_xuKHyEYNWcM_W000022 · in -14.4 dBFS · gain -5.6 dB · emolia-02260
(distress, sadness, concentration· steady, no disfluency, light breath, newsreading)Since 2011, Yemen has been in a state of political crisis starting with street protests against poverty, unemployment, corruption, and President Saleh's plan to amend Yemen's constitution and eliminate the presidential term limit, in effect making him president for life.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as distress, sadness, concentration; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 15.2s, EN.
EN_xuKHyEYNWcM_W000023 · in -15.3 dBFS · gain -4.7 dB · emolia-02260
(emotional numbness·fairly steady, no disfluency, light breath, newsreading)President Saleh stepped down and the powers of the presidency were transferred to Vice President Abdrabba Mansur Hadi, who was formally elected president on 21 February 2012 in a one-man election
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.7s, EN.
EN_xuKHyEYNWcM_W000024 · in -15.3 dBFS · gain -4.7 dB · emolia-02260
This chain comes from the two-sided rule: it only counts if both emotions move — Contempt down and Infatuation up — by at least 0.20 each.
The chain starts with Infatuation around average — 0.48, lower than 52 % of clips in this corpus — and ends with it strongly present at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.29.
At the same time Contempt goes the other way, from 0.92 (higher than 92 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.39. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.14 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.80 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.80. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 25 s · zh · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.764 before conversion and 0.666 after — it fell by 0.098. Neighbour-to-neighbour the worst pair went 0.707 → 0.666. (The earlier render, with segment 1 left raw, scores 0.599 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.292 in the original and +0.568 after conversion — 195 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contempt, -0.391 became +0.000.
Quality. Mean predicted overall quality across the segments went 2.80 → 2.98 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.764 → 0.666-0.098identity cos neighbours 0.707 → 0.666d_b rescored +0.292 → +0.568d_a rescored -0.391 → +0.000d_a mined -0.391d_b mined 0.292min_cos_consec (site) 0.8572min_cos_anchor (site) 0.8006dataset emolialang zhspeaker ZH_B00004_S04854total 24.1schain gain +2.2 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, normally alert, slightly relaxed, light breath
(contempt · normal-paced, fairly steady, no disfluency, authoritative)你会习惯于这种挣钱的模式。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contempt; style: authoritative, casual; average recording, no background noise; genuineness 2.6/6; vocal-burst blend 1.8/10; 3.1s, ZH.
ZH_B00004_S04854_W000011 · in -19.0 dBFS · gain -1.0 dB · emolia-03316
(pride, sourness·measured, fairly steady, some disfluency, monologue)如果你比较喜欢像自由职业者那样工作,你也很难跳出这种挣钱模式。如果你习惯了政府救济,这仍然是一个难以打破的形式。所以说从左象限进入右象限,最难的一关就是。
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, sourness; style: monologue, didactic; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 2.8/10; 18.0s, ZH.
ZH_B00004_S04854_W000012 · in -17.9 dBFS · gain -2.1 dB · emolia-03316
(measured, steady, no disfluency, formal)你必须改变目前的挣钱方式。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, didactic; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 1.5/10; 3.5s, ZH.
ZH_B00004_S04854_W000013 · in -20.1 dBFS · gain +0.1 dB · emolia-03316
This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Anger up — by at least 0.20 each.
The chain starts with Anger below average — 0.28, lower than 72 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.57.
At the same time Contemplation goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.45 (lower than 55 % of clips in this corpus), a change of -0.48. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.22, then +0.16 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.91 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 48 s · zh · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.727 before conversion and 0.705 after — it fell by 0.022. Neighbour-to-neighbour the worst pair went 0.792 → 0.800. (The earlier render, with segment 1 left raw, scores 0.625 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Anger moved +0.574 in the original and +0.351 after conversion — 61 % of the delta retained. On the other named axis, Contemplation, -0.479 became -0.367.
Quality. Mean predicted overall quality across the segments went 3.07 → 3.26 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.727 → 0.705-0.022identity cos neighbours 0.792 → 0.800d_b rescored +0.574 → +0.351d_a rescored -0.479 → -0.367d_a mined -0.479d_b mined 0.573min_cos_consec (site) 0.9104min_cos_anchor (site) 0.9212dataset emolialang zhspeaker ZH_B00000_S03792total 46.6schain gain +1.9 dBseam step 0.7 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a child masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(contemplation, interest, jealousy and envy · normal-paced, some disfluency, average clarity, monologue)北上资金呢被人们称作是聪明的资金,或者说外资啊,通常被人们称为是聪明的资金。在国内呢很多时候外资呢指的就是北上资金。其实呢外资它是一个统称,指的就是境外的投资者,境外的投资者看好a股市场。那么买进来的这个钱就叫外资,外资呢也包含了北上资金,因为香港地区呢属于是境外,而且香港股市是一个开放型的市场。
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, interest, jealousy and envy; style: monologue, storytelling; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 7.4/10; 27.6s, ZH.
ZH_B00000_S03792_W000001 · in -20.2 dBFS · gain +0.2 dB · emolia-03267
(fast, some disfluency, clear, authoritative)除了当地的投资者之外,也有很多海外投资者在香港投资。那么自从开通了沪深港通之后。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, monologue; average recording, no background noise; genuineness 1.8/6; vocal-burst blend 4.5/10; 7.2s, ZH.
ZH_B00000_S03792_W000002 · in -20.2 dBFS · gain +0.2 dB · emolia-03267
(infatuation·normal-paced, no disfluency, clear, authoritative)海外的投资者可以在香港通过沪深港通来买卖a股。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation; style: authoritative, formal; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 2.5/10; 4.5s, ZH.
ZH_B00000_S03792_W000003 · in -19.8 dBFS · gain -0.2 dB · emolia-03267
(fast, some disfluency, clear, authoritative)所以呢当北上资金大幅流入a股市场的时候,它不一定就是香港投资者在买,也可能是海外投资者在买。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, didactic; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 4.1/10; 7.9s, ZH.
ZH_B00000_S03792_W000004 · in -21.1 dBFS · gain +1.1 dB · emolia-03267