PXR at chain length k=5, all corpora, at the mining floor.
This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_k-PXR-k5.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words
are a real non-speech sound, printed where it happens.
The full generated caption for any clip is under “full caption & clip
details”. Its perceived-gender and background-noise clauses were re-rendered from
the numeric buckets, because the versions stored in the corpus index had those two
ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
The tags and captions describe the ORIGINAL clips — they are the
corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the
converted audio are the ones printed in each card's own paragraph and chips.
This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Contemplation clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.30.
At the same time Contempt goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.15, then -0.19, then +0.20, then +0.14 — not a clean run: step 2 moves back the other way by 0.19 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.37 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.38 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.37, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 77 s · pt · eurospeech
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.357 before conversion and 0.854 after — it rose by 0.496. Neighbour-to-neighbour the worst pair went 0.370 → 0.838. (The earlier render, with segment 1 left raw, scores 0.796 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.299 in the original and +0.308 after conversion — 103 % of the delta retained, which is essentially all of it. On the other named axis, Contempt, -0.319 became -0.094.
Quality. Mean predicted overall quality across the segments went 3.23 → 3.38 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.357 → 0.854+0.496identity cos neighbours 0.370 → 0.838d_b rescored +0.299 → +0.308d_a rescored -0.319 → -0.094d_a mined -0.319d_b mined 0.299min_cos_consec (site) 0.3840min_cos_anchor (site) 0.3702dataset eurospeechlang ptspeaker portugal_16_1_15total 76.1schain gain +3.5 dBseam step 1.3 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an elderly somewhat feminine voice · quiet background, slightly relaxed, frequent disfluency, fairly narrow pitch
(contempt · measured, normally alert, fairly steady, monologue)O jovem atleta lusodescendente, com raízes em Braga, foi formado no Clermont e jogava atualmente no Vienne Rugby. A sua trágica partida deixa uma enorme lacuna no mundo do rugby
full caption & clip details
An elderly somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contempt; style: monologue, narration; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 2.6/10; 14.8s, PT.
portugal_16_1_15_10039008_10053776 · in -27.5 dBFS · gain +7.5 dB · eurospeech-02639
(contempt, disappointment, malevolence malice· measured, very low-energy, fairly steady, monologue)será sempre recordado pelo seu talento, determinação, dedicação e paixão pelo desporto. Perante a perda precoce deste jovem atleta, a Assembleia da República manifesta o seu mais profundo pesar pelo seu falecimento,
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, rough, thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, neutral stance, slightly guarded; reads as contempt, disappointment, malevolence malice; style: monologue, whispered; poor recording, quiet background; genuineness 2.9/6; vocal-burst blend 2.0/10; 14.6s, PT.
portugal_16_1_15_10053776_10068367 · in -28.4 dBFS · gain +8.3 dB · eurospeech-02639
(concentration, malevolence malice ·slow, subdued, steady, monologue)Além da capital, Porto Alegre, 417 dos 497 municípios da região foram afetados, com destaque para os vales dos rios Taquari, Caí, Pardo, Jacuí, Sinos e Gravataí.
full caption & clip details
A child masculine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, malevolence malice; style: monologue, authoritative; average recording, quiet background; genuineness 0.5/6; vocal-burst blend 1.1/10; 17.7s, PT.
portugal_16_1_15_10104654_10122351 · in -30.5 dBFS · gain +10.5 dB · eurospeech-02639
(triumph, pride, malevolence malice · slow, subdued, steady, cartoonish)O número de vítimas mortais tem crescido todos os dias e ultrapassa já uma centena. Há ainda vários feridos, desaparecidos e desalojados. As autoridades estimam que o número total de afetados é já de 1 milhão e 400 mil pessoas.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, rough, balanced body; clear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, fairly guarded; reads as triumph, pride, malevolence malice; style: cartoonish, monologue; average recording, quiet background; genuineness 0.7/6; vocal-burst blend 0.6/10; 17.3s, PT.
portugal_16_1_15_10122351_10139664 · in -28.9 dBFS · gain +8.9 dB · eurospeech-02639
(contemplation, emotional numbness· slow, normally alert, steady, monologue)Os temporais continuam a fazer-se sentir, o que é motivo de preocupação e pode levar a um drama ainda maior. Neste sentido, o Presidente da Assembleia da República
full caption & clip details
A child masculine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, fairly guarded; reads as contemplation, emotional numbness; style: monologue, authoritative; average recording, quiet background; genuineness 0.1/6; vocal-burst blend 1.0/10; 12.5s, PT.
portugal_16_1_15_10139664_10152160 · in -28.5 dBFS · gain +8.5 dB · eurospeech-02639
This chain comes from the proxy rule: the same two-sided test as above, but because Jealousy and Envy is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Jealousy and Envy around average — 0.58, higher than 58 % of clips in this corpus — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.30.
At the same time Concentration goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.24, then -0.17, then +0.09, then +0.14 — not a clean run: step 2 moves back the other way by 0.17 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 71 s · da · eurospeech
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.924 before conversion and 0.909 after — it fell by 0.015. Neighbour-to-neighbour the worst pair went 0.924 → 0.909. (The earlier render, with segment 1 left raw, scores 0.776 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Jealousy and Envy moved +0.291 in the original and +0.145 after conversion — 50 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Concentration, -0.250 became -0.385.
Quality. Mean predicted overall quality across the segments went 3.07 → 3.29 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.924 → 0.909-0.015identity cos neighbours 0.924 → 0.909d_b rescored +0.291 → +0.145d_a rescored -0.250 → -0.385d_a mined -0.251d_b mined 0.299min_cos_consec (site) 0.9389min_cos_anchor (site) 0.9528dataset eurospeechlang daspeaker denmark_20181M004_2018-10-total 69.5schain gain +2.4 dBseam step 0.5 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a middle-aged feminine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, normally alert, slightly relaxed, fairly steady
(concentration, malevolence malice · normal-paced, little disfluency, light breath, monologue)det handler om forskning. Det er mange af de samme ting, som det handler om for befolkningen i resten af verden. Derfor er der også en vigtig pointe for Inuit Ataqatigiit at sige, at vi skal have fokus på den menneskelige dimension i Arktis.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, malevolence malice; style: monologue, formal; average recording, quiet background; genuineness 0.6/6; vocal-burst blend 0.5/10; 16.4s, DA.
denmark_20181M004_2018-10-09_1300_20917199_20933568 · in -22.0 dBFS · gain +2.0 dB · eurospeech-00335
(disgust, malevolence malice ·measured, little disfluency, audible breath, monologue)Vi er ikke særlig mange mennesker i Arktis i forhold til i den sydligere del af verden, og derfor tæller hver en menneskelig ressource i Arktis.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, little disfluency, moderate pitch range, audible breath; affect is mildly positive, neutral stance, slightly guarded; reads as disgust, malevolence malice; style: monologue, authoritative; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 1.0/10; 10.4s, DA.
denmark_20181M004_2018-10-09_1300_20933568_20943929 · in -21.8 dBFS · gain +1.8 dB · eurospeech-00335
(disgust, concentration·normal-paced, some disfluency, light breath, monologue)Som Folketingets repræsentant i komiteen for arktiske parlamentarikere oplever jeg heldigvis også, at vi snakker meget mere om erhvervsudvikling, jobskabelse og øget digitalisering i Arktis.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as disgust, concentration; style: monologue, didactic; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 0.8/10; 13.0s, DA.
denmark_20181M004_2018-10-09_1300_20943929_20956880 · in -22.3 dBFS · gain +2.3 dB · eurospeech-00335
(concentration · normal-paced, frequent disfluency, light breath, monologue)I Grønland har 83 pct. af befolkningen (low mumble) adgang til internettet, og ud af de 56.000 mennesker, som vi er i Grønland, er 43.000 mennesker på Facebook – bare for at komme med nogle eksempler.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.8/10; 14.6s, DA.
denmark_20181M004_2018-10-09_1300_20956880_20971520 · in -21.8 dBFS · gain +1.8 dB · eurospeech-00335
(measured, some disfluency, light breath, didactic)Grønland er også et moderne samfund, ligesom det også er et meget traditionelt samfund. Med de få mennesker, som vi har i Arktis, kan vi ikke kun tænke kommercielt. Vi er også nødt til at tænke kreativt og alternativt.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 0.7/10; 15.9s, DA.
denmark_20181M004_2018-10-09_1300_20971520_20987424 · in -21.9 dBFS · gain +1.9 dB · eurospeech-00335
This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Interest below average — 0.38, lower than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.55.
At the same time Distress goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.44 (lower than 56 % of clips in this corpus), a change of -0.52. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.10, then +0.17, then +0.07 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.76 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.76 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.76, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 51 s · en · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.686 before conversion and 0.734 after — it rose by 0.048. Neighbour-to-neighbour the worst pair went 0.804 → 0.844. (The earlier render, with segment 1 left raw, scores 0.630 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.553 in the original and +0.535 after conversion — 97 % of the delta retained, which is essentially all of it. On the other named axis, Distress, -0.524 became -0.533.
Quality. Mean predicted overall quality across the segments went 2.89 → 3.13 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.686 → 0.734+0.048identity cos neighbours 0.804 → 0.844d_b rescored +0.553 → +0.535d_a rescored -0.524 → -0.533d_a mined -0.523d_b mined 0.547min_cos_consec (site) 0.7612min_cos_anchor (site) 0.7612dataset podcastlang enspeaker 270299total 49.8schain gain +3.4 dBseam step 1.2 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, good recording, normally alert, average clarity, moderate pitch range, light breath
(distress, disappointment, impatience and irritability · brisk, slightly relaxed, fairly steady, narration)And if there isn't a massive joy of recycling for you, why are you sharing it with the children?
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as distress, disappointment, impatience and irritability; style: narration, storytelling; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 1.5/10; 5.0s, EN.
270299_00327652 · in -30.2 dBFS · gain +10.2 dB · podcast-04891
(normal-paced, slightly relaxed, moderately variable, playful)It's it's fruitless. It's like, you know, it's why drawing clubs got no set texts.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: playful, casual; good recording, no background noise; genuineness 3.3/6; vocal-burst blend 2.3/10; 5.0s, EN.
270299_00329504 · in -31.4 dBFS · gain +11.4 dB · podcast-04889
(sourness, teasing, embarrassment· normal-paced, neutral tension, moderately variable, casual)There aren't any. What's your joy? You know, it's like there's certain books I really don't like sharing with children 'cause I actually find them incredibly boring.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as sourness, teasing, embarrassment; style: casual, whispered; good recording, quiet background; genuineness 4.1/6; vocal-burst blend 4.8/10; 9.8s, EN.
270299_00329996 · in -30.1 dBFS · gain +10.1 dB · podcast-04886
(contemplation, doubt, embarrassment · normal-paced, neutral tension, moderately variable, conversational)(ahem) Um so why would I you know if I I might share them at home time, but I'm not gonna make a you know, I'm not gonna make a massive song and a dance about them. There might be books I think, do you know what these children will enjoy? So of course I'm gonna share it with them. But the only way I'm gonna find that out is if I go and play with them.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as contemplation, doubt, embarrassment; style: conversational, casual; good recording, quiet background; genuineness 4.3/6; vocal-burst blend 8.9/10; 15.6s, EN.
270299_00330972 · in -30.4 dBFS · gain +10.4 dB · podcast-04886
(interest, pride, jealousy and envy·brisk, slightly relaxed, fairly steady, monologue)Do you know what I mean? It's like I'm not necessarily switched on by books about tractors, for example. But if I've got loads of children that are really fascinated by tractors, then because I'm the companion, I'm gonna go away and I'm gonna discover the joy of tractors, and now lo and behold, here's a book about tractors.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, pride, jealousy and envy; style: monologue, formal; good recording, quiet background; genuineness 2.1/6; vocal-burst blend 3.1/10; 15.2s, EN.
270299_00332524 · in -29.4 dBFS · gain +9.4 dB · podcast-04887
This chain comes from the proxy rule: the same two-sided test as above, but because Fear is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Fear around average — 0.56, higher than 56 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.28.
At the same time Concentration goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.14, then -0.21, then +0.17, then +0.18 — not a clean run: step 2 moves back the other way by 0.21 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.79 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 39 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.785 before conversion and 0.780 after — it fell by 0.005. Neighbour-to-neighbour the worst pair went 0.797 → 0.763. (The earlier render, with segment 1 left raw, scores 0.690 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.282 in the original and +0.125 after conversion — 44 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Concentration, -0.270 became -0.241.
Quality. Mean predicted overall quality across the segments went 2.50 → 2.91 (+0.41) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.785 → 0.780-0.005identity cos neighbours 0.797 → 0.763d_b rescored +0.282 → +0.125d_a rescored -0.270 → -0.241d_a mined -0.272d_b mined 0.281min_cos_consec (site) 0.7931min_cos_anchor (site) 0.8445dataset emolialang enspeaker EN_y2jsA1HNOYQtotal 37.7schain gain +3.7 dBseam step 1.3 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, clear, moderate pitch range
(concentration · steady, almost no disfluency, formal, monologue)This patent intended to achieve this state through fabrication of metal and ceramic implantable articulation members with small channels throughout the design.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration; style: formal, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.1/10; 8.3s, EN.
EN_y2jsA1HNOYQ_W000065 · in -15.8 dBFS · gain -4.2 dB · emolia-01296
(concentration ·fairly steady, almost no disfluency, monologue, formal)Allowing synovial fluid, the hip capsule's natural lubricant, to move through the articulation member and create the desired excess lubrication layer necessary to mimic natural hip joint mechanics.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration; style: monologue, formal; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.8/10; 11.4s, EN.
EN_y2jsA1HNOYQ_W000066 · in -15.3 dBFS · gain -4.7 dB · emolia-01296
(concentration · fairly steady, some disfluency, monologue, formal)This design utilizes surface engineering by manipulating roughness in the form of micro-dimpling on the surface of the prosthesis to better maintain the lubrication layer.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration; style: monologue, formal; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 0.5/10; 9.4s, EN.
EN_y2jsA1HNOYQ_W000068 · in -14.9 dBFS · gain -5.1 dB · emolia-01296
(emotional numbness· fairly steady, almost no disfluency, formal, authoritative)Microdimples act as reservoirs for the lubricant in the prosthesis, which maintain an increased lubrication.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness; style: formal, authoritative; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.2/10; 5.8s, EN.
EN_y2jsA1HNOYQ_W000069 · in -14.5 dBFS · gain -5.5 dB · emolia-01296
(fairly steady, no disfluency, formal, casual)It should be noted that a lubricant must be injected into the prosthetic joint
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, casual; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 1.4/10; 3.5s, EN.
EN_y2jsA1HNOYQ_W000070 · in -15.3 dBFS · gain -4.7 dB · emolia-01296
This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Emotional Numbness clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.27.
At the same time Hope Enthusiasm Optimism goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.50 (right about the corpus median), a change of -0.40. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.14, then -0.14, then +0.23, then +0.04 — not a clean run: step 2 moves back the other way by 0.14 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 58 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.803 before conversion and 0.844 after — it rose by 0.040. Neighbour-to-neighbour the worst pair went 0.803 → 0.761. (The earlier render, with segment 1 left raw, scores 0.770 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.273 in the original and +0.339 after conversion — 124 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Hope Enthusiasm Optimism, -0.403 became -0.357.
Quality. Mean predicted overall quality across the segments went 2.98 → 3.11 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.803 → 0.844+0.040identity cos neighbours 0.803 → 0.761d_b rescored +0.273 → +0.339d_a rescored -0.403 → -0.357d_a mined -0.403d_b mined 0.273min_cos_consec (site) 0.9153min_cos_anchor (site) 0.9289dataset emolialang enspeaker EN_HrjTLDt0ZQutotal 56.4schain gain +1.5 dBseam step 3.7 dBcrossfades 100/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(hope enthusiasm optimism · normal-paced, no disfluency, light breath, formal)Facebook ad headline feature. Generate scroll stopping headlines for your Facebook ads to get prospects to click and ultimately buy.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.9/10; 7.2s, EN.
EN_HrjTLDt0ZQu_W000020 · in -19.6 dBFS · gain -0.4 dB · emolia-02313
(hope enthusiasm optimism · normal-paced, little disfluency, minimal breath, monologue)Par Framework feature, problem agitate solution. A valuable framework for creating new marketing copy ideas. Blog post topic ideas feature, brainstorms new blog post topics that will engage readers and rank well on Google. Aida framework feature, uses the oldest marketing framework in the world.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as hope enthusiasm optimism; style: monologue, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.6/10; 18.8s, EN.
EN_HrjTLDt0ZQu_W000021 · in -19.2 dBFS · gain -0.8 dB · emolia-02313
(contentment, hope enthusiasm optimism · normal-paced, almost no disfluency, light breath, newsreading)Content improve a feature. Takes a piece of content and rewrite it to make it more interesting, creative, and engaging. Sentence expand a feature. Expands a short sentence or a few words into a longer sentence that is creative, interesting, and engaging.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contentment, hope enthusiasm optimism; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.3/10; 14.9s, EN.
EN_HrjTLDt0ZQu_W000023 · in -19.5 dBFS · gain -0.5 dB · emolia-02313
(brisk, no disfluency, light breath, formal)Problem agitate solution feature is a valuable framework for creating new marketing copy ideas.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.9/10; 5.4s, EN.
EN_HrjTLDt0ZQu_W000024 · in -18.4 dBFS · gain -1.6 dB · emolia-02313
(normal-paced, no disfluency, light breath, formal)Product description feature, creates compelling product descriptions to be used on websites, emails and social media. Blog post outline feature, creates lists and outlines for articles.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.9s, EN.
EN_HrjTLDt0ZQu_W000025 · in -19.4 dBFS · gain -0.6 dB · emolia-02313
This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Contemplation clearly present — 0.58, higher than 58 % of clips in this corpus — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.29.
At the same time Shame goes the other way, from 0.88 (higher than 88 % of clips in this corpus) to 0.58 (higher than 58 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are -0.11, then -0.05, then +0.22, then +0.23 — not a clean run: step 1 moves back the other way by 0.11 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.85 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 51 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.869 before conversion and 0.836 after — it fell by 0.033. Neighbour-to-neighbour the worst pair went 0.713 → 0.686. (The earlier render, with segment 1 left raw, scores 0.775 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.295 in the original and +0.279 after conversion — 95 % of the delta retained, which is essentially all of it. On the other named axis, Shame, -0.304 became -0.300.
Quality. Mean predicted overall quality across the segments went 3.02 → 3.18 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.869 → 0.836-0.033identity cos neighbours 0.713 → 0.686d_b rescored +0.295 → +0.279d_a rescored -0.304 → -0.300d_a mined -0.304d_b mined 0.294min_cos_consec (site) 0.8528min_cos_anchor (site) 0.9173dataset emolialang zhspeaker ZH_B00041_S08342total 49.9schain gain +1.5 dBseam step 2.2 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady, no disfluency
This chain comes from the proxy rule: the same two-sided test as above, but because Hope Enthusiasm Optimism is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Hope Enthusiasm Optimism clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.26.
At the same time Doubt goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are -0.02, then +0.10, then +0.03, then +0.15 — not a clean run: step 1 moves back the other way by 0.02 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.23 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.23 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.23, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 80 s · en · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.234 before conversion and 0.662 after — it rose by 0.428. Neighbour-to-neighbour the worst pair went 0.220 → 0.681. (The earlier render, with segment 1 left raw, scores 0.468 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.260 in the original and +0.483 after conversion — 186 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Doubt, -0.257 became -0.283.
Quality. Mean predicted overall quality across the segments went 3.03 → 3.19 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.234 → 0.662+0.428identity cos neighbours 0.220 → 0.681d_b rescored +0.260 → +0.483d_a rescored -0.257 → -0.283d_a mined -0.257d_b mined 0.259min_cos_consec (site) 0.2261min_cos_anchor (site) 0.2340dataset podcastlang enspeaker 586305total 78.5schain gain +4.6 dBseam step 2.3 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, fairly steady
(doubt, contemplation, relief · normally alert, neutral tension, conversational, casual)(low mumble) Um, you know, not not that old. Uh (low mumble) some of them are probably older, but I I think a lot of them have been built and employed within the last decade or so. Okay.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, contemplation, relief; style: conversational, casual; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 4.0/10; 9.4s, EN.
586305_00118504 · in -21.9 dBFS · gain +1.9 dB · podcast-04364
(interest, fear, contemplation · normally alert, neutral tension, casual, conversational)I mean, is there like an ancient Chinese tradition of having like (ahem) uh fishing boats as part of your military or something? Because that would (ahem) uh or was it like this like doesn't seem like a more conscious strategy to just invent something new. I
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, fear, contemplation; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 10.0/10; 11.2s, EN.
586305_00119463 · in -21.8 dBFS · gain +1.8 dB · podcast-04368
(doubt, interest, awe· normally alert, slightly relaxed, casual, conversational)don't know if there's an ancient Chinese tradition of it. Uh (low mumble) but this particular these particular gray zone tactics, these are something that we've really seen China use a lot of within the last decade, I would say. Okay.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as doubt, interest, awe; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 1.4/10; 12.6s, EN.
586305_00120576 · in -22.5 dBFS · gain +2.5 dB · podcast-04380
(concentration, jealousy and envy, contemplation·subdued, slightly relaxed, conversational, casual)So what about okay? So you have you would you would invade okay, so you you have the capabilities to invade it's tough because you know they're they know where the Taiwanese know where the sites are. Uh (ahem) what about what about air power? What what is what what can (low mumble) uh what can China do? We've been watching this in the Russia war. I mean, I was I I you know, I'm surprised how Little Russia could do in the infrastructure. I thought having missiles and stuff, you could knock out a you know a country's power, but apparently that's not the easiest thing in the world to do. (low mumble) Um, you know, what can China do sort of from the air (ahem) uh to Taiwan?
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as concentration, jealousy and envy, contemplation; style: conversational, casual; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 7.2/10; 29.2s, EN.
586305_00121848 · in -22.3 dBFS · gain +2.3 dB · podcast-04391
(normally alert, slightly relaxed, monologue)Well, they can do a lot. (low mumble) Uh China's been developing its air force, (ahem) uh, you know, its its bombers, its, its fighters, (ahem) uh its aircraft carriers. And and so certainly it could contribute a lot of firepower from the air. Now, I want to be careful.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.6/10; 16.9s, EN.
586305_00124800 · in -22.7 dBFS · gain +2.7 dB · podcast-02762
This chain comes from the proxy rule: the same two-sided test as above, but because Hope Enthusiasm Optimism is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Hope Enthusiasm Optimism clearly present — 0.58, higher than 58 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.38.
At the same time Embarrassment goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.09, then +0.22, then +0.09, then -0.01 — not a clean run: step 4 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.08 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.08 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.08, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 57 s · en · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.271 before conversion and 0.342 after — it rose by 0.071. Neighbour-to-neighbour the worst pair went 0.176 → 0.369. (The earlier render, with segment 1 left raw, scores 0.285 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.381 in the original and +0.376 after conversion — 99 % of the delta retained, which is essentially all of it. On the other named axis, Embarrassment, -0.276 became -0.348.
Quality. Mean predicted overall quality across the segments went 2.69 → 3.01 (+0.32) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.271 → 0.342+0.071identity cos neighbours 0.176 → 0.369d_b rescored +0.381 → +0.376d_a rescored -0.276 → -0.348d_a mined -0.275d_b mined 0.381min_cos_consec (site) 0.0788min_cos_anchor (site) 0.0788dataset podcastlang enspeaker 865643total 55.2schain gain +3.5 dBseam step 1.1 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, quiet background, normal-paced, moderately variable, some disfluency, wide pitch range, light breath
(embarrassment, amusement, teasing · normally alert, neutral tension, average clarity, casual)Okay, yeah, I'm bad at these questions. That's what you all learned about me.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as embarrassment, amusement, teasing; style: casual, conversational; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 3.9/10; 3.3s, EN.
865643_00265576 · in -19.8 dBFS · gain -0.2 dB · podcast-03610
(affection, embarrassment, confusion· normally alert, neutral tension, average clarity, conversational)Actually, the reason the reason that I thought of that is I don't know where you guys are with the chosen, whatever, but (low mumble) um I I I would love to be friends with Barnaby.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as affection, embarrassment, confusion; style: conversational, casual; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 5.2/10; 10.5s, EN.
865643_00266344 · in -20.0 dBFS · gain +0.0 dB · podcast-03597
(affection, infatuation, pleasure ecstasy· normally alert, neutral tension, average clarity, casual)I his character is just it's comic relief in the series, but I just I absolutely every time he's on the screen, I'm like, I would love to be friends. (low mumble)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as affection, infatuation, pleasure ecstasy; style: casual, playful; average recording, quiet background; mildly explicit content; genuineness 5.6/6; vocal-burst blend 4.1/10; 13.9s, EN.
865643_00267696 · in -21.9 dBFS · gain +1.9 dB · podcast-01294
(thankfulness gratitude, contentment, affection ·very low-energy, relaxed, average clarity, conversational)Yeah. So ladies, thanks so much for taking some time to be a part of this. Amy, great answers. And I know you have encouraged me. And I guarantee you have encouraged others that are out there as well. Any last words you'd like to share with
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as thankfulness gratitude, contentment, affection; style: conversational, casual; average recording, quiet background; mildly explicit content; genuineness 4.4/6; vocal-burst blend 5.2/10; 18.6s, EN.
865643_00269080 · in -25.7 dBFS · gain +5.7 dB · podcast-01296
(hope enthusiasm optimism, doubt, thankfulness gratitude ·normally alert, neutral tension, somewhat unclear, casual)us? (ahem) I'd encourage you to come to the worship leader gathering if you're able in October. Maybe my commitment will be if you come, I'll have answers to those questions.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as hope enthusiasm optimism, doubt, thankfulness gratitude; style: casual, playful; below-average recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.1/10; 9.6s, EN.
865643_00270935 · in -23.3 dBFS · gain +3.3 dB · podcast-01308
This chain comes from the proxy rule: the same two-sided test as above, but because Confusion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Confusion clearly present — 0.58, higher than 58 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.26.
At the same time Emotional Numbness goes the other way, from 0.81 (higher than 81 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.01, then -0.13, then +0.15 — not a clean run: step 3 moves back the other way by 0.13 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.76 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.76 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.76, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 34 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.707 before conversion and 0.659 after — it fell by 0.048. Neighbour-to-neighbour the worst pair went 0.707 → 0.693. (The earlier render, with segment 1 left raw, scores 0.640 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.259 in the original and +0.767 after conversion — 296 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, -0.273 became -0.265.
Quality. Mean predicted overall quality across the segments went 2.92 → 3.13 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.707 → 0.659-0.048identity cos neighbours 0.707 → 0.693d_b rescored +0.259 → +0.767d_a rescored -0.273 → -0.265d_a mined -0.273d_b mined 0.259min_cos_consec (site) 0.7568min_cos_anchor (site) 0.7568dataset emolialang zhspeaker ZH_B00022_S00443total 32.6schain gain +0.1 dBseam step 1.6 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · fairly smooth, balanced body, no background noise, slightly relaxed, fairly steady, light breath
(measured, normally alert, some disfluency, authoritative)鸭子的话语却几乎摧毁了我心底所有的防线,令我的鼻子又酸又涩。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, monologue; average recording, no background noise; genuineness 2.6/6; vocal-burst blend 4.5/10; 6.3s, ZH.
ZH_B00022_S00443_W000015 · in -20.0 dBFS · gain -0.0 dB · emolia-03495
(fear, emotional numbness, distress· measured, subdued, little disfluency, narration)你明不明白,这是天意,我和雷震子不熟,不过老三,他肯定是一个义薄云天的好兄弟。
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is warm, slightly dark, fairly smooth, balanced body; somewhat unclear, little disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, fairly guarded; reads as fear, emotional numbness, distress; style: narration, monologue; average recording, no background noise; genuineness 2.0/6; vocal-burst blend 5.6/10; 8.6s, ZH.
ZH_B00022_S00443_W000016 · in -21.9 dBFS · gain +1.9 dB · emolia-03495
(pride, contemplation, shame· measured, normally alert, some disfluency, didactic)因为雷震子用他的命救了你们的命,而我留在这里,就是为了确保雷震子的一条命没有白费。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, contemplation, shame; style: didactic, authoritative; average recording, no background noise; genuineness 1.9/6; vocal-burst blend 2.8/10; 7.9s, ZH.
ZH_B00022_S00443_W000017 · in -18.5 dBFS · gain -1.5 dB · emolia-03495
(impatience and irritability·normal-paced, normally alert, no disfluency, formal)就是为了保证你们这几天之内哪里都不能去。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as impatience and irritability; style: formal, authoritative; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 3.7/10; 3.3s, ZH.
ZH_B00022_S00443_W000018 · in -20.4 dBFS · gain +0.4 dB · emolia-03495
(measured, normally alert, little disfluency, narration)心脏在胸腔内剧烈的跳动声如同雷鸣一般,在我的耳廓深处响起,我痴痴的望着鸭子。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 3.8/10; 7.3s, ZH.
ZH_B00022_S00443_W000019 · in -21.3 dBFS · gain +1.3 dB · emolia-03495
This chain comes from the proxy rule: the same two-sided test as above, but because Relief is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Relief barely there — 0.23, lower than 77 % of clips in this corpus — and ends with it clearly present at 0.65, higher than 65 % of clips in this corpus. That is a total rise of 0.42.
At the same time Doubt goes the other way, from 0.87 (higher than 87 % of clips in this corpus) to 0.49 (lower than 51 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.19, then -0.02, then +0.18, then +0.07 — not a clean run: step 2 moves back the other way by 0.02 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 54 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.752 before conversion and 0.807 after — it rose by 0.055. Neighbour-to-neighbour the worst pair went 0.712 → 0.789. (The earlier render, with segment 1 left raw, scores 0.779 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.424 in the original and +0.410 after conversion — 97 % of the delta retained, which is essentially all of it. On the other named axis, Doubt, -0.380 became -0.027.
Quality. Mean predicted overall quality across the segments went 2.92 → 3.23 (+0.31) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.752 → 0.807+0.055identity cos neighbours 0.712 → 0.789d_b rescored +0.424 → +0.410d_a rescored -0.380 → -0.027d_a mined -0.380d_b mined 0.424min_cos_consec (site) 0.8854min_cos_anchor (site) 0.8642dataset emolialang zhspeaker ZH_B00034_S01581total 52.4schain gain +3.3 dBseam step 1.8 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(normal-paced, no disfluency, clear, monologue)就像是那种悬疑小说里的暗格一样,叶宵打开一看,哎,里面果然装着一个盒子,盒子里面是一块牙齿,估计是妈妈的遗物,所以呢也残留着能量。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 3.7/10; 10.2s, ZH.
ZH_B00034_S01581_W000006 · in -18.2 dBFS · gain -1.8 dB · emolia-03613
(fast, no disfluency, clear, authoritative)叶相用小锤子咔嚓一敲牙齿被敲碎,而永子的诅咒印记终于也随之消失。
full caption & clip details
An adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; average recording, no background noise; genuineness 0.7/6; vocal-burst blend 4.2/10; 6.2s, ZH.
ZH_B00034_S01581_W000007 · in -17.2 dBFS · gain -2.8 dB · emolia-03613
(interest, disappointment·brisk, some disfluency, average clarity, monologue)倒在地上的父亲不停的忏悔,他犯下了太多错误,而且还把十七个无辜的人给害死了。看到这一幕的叶宵,决定留下父亲的灵,让他留在这个灵地,继续当屋主忏悔,替夜宵管理将来的毕业生之家,赶走那些可能不小心进入的活人。
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, disappointment; style: monologue, authoritative; good recording, quiet background; genuineness 1.3/6; vocal-burst blend 7.6/10; 16.3s, ZH.
ZH_B00034_S01581_W000008 · in -17.2 dBFS · gain -2.8 dB · emolia-03613
(interest · brisk, almost no disfluency, average clarity, monologue)永子也被父亲的忏悔打动,流着泪打电话给卖家,决定买下这座房子。而一旁的夜宵,看到自己的骷髅恶灵若有所思的样子,难难道你的记忆和力量已经稍微恢复了吗?看来这个骷髅恶灵啊也有着非同一般的来头,所以才会被叶宵一直带在身边。
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest; style: monologue, authoritative; good recording, quiet background; genuineness 1.2/6; vocal-burst blend 7.1/10; 16.4s, ZH.
ZH_B00034_S01581_W000009 · in -17.4 dBFS · gain -2.6 dB · emolia-03613
(fast, no disfluency, clear, authoritative)终于,凶宅之旅结束了毕业生之家get to does z.
full caption & clip details
An adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 2.9/10; 4.0s, ZH.
ZH_B00034_S01581_W000010 · in -16.0 dBFS · gain -4.0 dB · emolia-03613
This chain comes from the proxy rule: the same two-sided test as above, but because Fear is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Fear clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.30.
At the same time Contemplation goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.46 (lower than 54 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are -0.03, then -0.15, then +0.23, then +0.25 — not a clean run: step 1 moves back the other way by 0.03 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 35 s · ko · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.815 before conversion and 0.831 after — it rose by 0.015. Neighbour-to-neighbour the worst pair went 0.775 → 0.831. (The earlier render, with segment 1 left raw, scores 0.738 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.298 in the original and +0.250 after conversion — 84 % of the delta retained, which is most of it. On the other named axis, Contemplation, -0.404 became -0.243.
Quality. Mean predicted overall quality across the segments went 2.98 → 3.14 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.815 → 0.831+0.015identity cos neighbours 0.775 → 0.831d_b rescored +0.298 → +0.250d_a rescored -0.404 → -0.243d_a mined -0.404d_b mined 0.298min_cos_consec (site) 0.8573min_cos_anchor (site) 0.8214dataset emolialang kospeaker KO_5ArkEucqmE4total 33.4schain gain +1.9 dBseam step 0.2 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, light breath
(normal-paced, fairly steady, little disfluency, formal)제가 최근에 읽은 책 가운데 마음에 감동이 있는 내용을 구독자 분들과 함께 나누는 시간인데요.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.4/10; 6.4s, KO.
KO_5ArkEucqmE4_W000001 · in -18.1 dBFS · gain -1.9 dB · emolia-03185
(intoxication altered states of consciousness, longing·measured, steady, no disfluency, formal)(ahem) 오늘 책은 지난번에 이어서 양순자님의 어른 공부입니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness, longing; style: formal, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.3/10; 6.2s, KO.
KO_5ArkEucqmE4_W000002 · in -18.8 dBFS · gain -1.2 dB · emolia-03185
(measured, fairly steady, some disfluency, formal)말씀드렸지만, 양순자님은 돌아가셨습니다. 30년 동안.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.3/10; 5.2s, KO.
KO_5ArkEucqmE4_W000003 · in -17.5 dBFS · gain -2.5 dB · emolia-03185
(measured, fairly steady, some disfluency, formal)교도소 사형수들을 상담하는 종교위원으로 봉사하시다가,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 0.8/10; 4.8s, KO.
KO_5ArkEucqmE4_W000004 · in -18.0 dBFS · gain -2.0 dB · emolia-03185
(fear· measured, steady, little disfluency, didactic)이 책이 출판되고 돌아가셨습니다. 그런데 이 책은 많은 독자들에게 잔잔한 감동을 주었고, 책은 다시 개정판이 나왔습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, little disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear; style: didactic, monologue; average recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.1/10; 11.6s, KO.
KO_5ArkEucqmE4_W000005 · in -18.9 dBFS · gain -1.1 dB · emolia-03185
This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Concentration around average — 0.49, lower than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.49.
At the same time Fatigue Exhaustion goes the other way, from 0.73 (higher than 73 % of clips in this corpus) to 0.31 (lower than 69 % of clips in this corpus), a change of -0.42. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are -0.01, then +0.17, then +0.22, then +0.11 — not a clean run: step 1 moves back the other way by 0.01 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 36 s · de · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.758 before conversion and 0.752 after — it fell by 0.006. Neighbour-to-neighbour the worst pair went 0.835 → 0.738. (The earlier render, with segment 1 left raw, scores 0.661 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.486 in the original and +0.379 after conversion — 78 % of the delta retained, which is most of it. On the other named axis, Fatigue Exhaustion, -0.417 became -0.086.
Quality. Mean predicted overall quality across the segments went 2.95 → 3.09 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.758 → 0.752-0.006identity cos neighbours 0.835 → 0.738d_b rescored +0.486 → +0.379d_a rescored -0.417 → -0.086d_a mined -0.417d_b mined 0.486min_cos_consec (site) 0.8710min_cos_anchor (site) 0.8710dataset emolialang despeaker DE_1UW85J2KzTutotal 34.9schain gain +1.0 dBseam step 0.9 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a child feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, slightly relaxed
(normally alert, fairly steady, formal, storytelling)ist ein Programm zum Sammeln und Verwalten von Literatur.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: formal, storytelling; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 3.6s, DE.
DE_1UW85J2KzTu_W000000 · in -20.9 dBFS · gain +0.9 dB · emolia-00167
(astonishment surprise·energised, fairly steady, storytelling, narration)Es ermöglicht das Verlinken von Literatur sowie das Zitieren dieser in verschiedenen Programmen.
full caption & clip details
A child feminine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as astonishment surprise; style: storytelling, narration; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 1.3/10; 5.9s, DE.
DE_1UW85J2KzTu_W000001 · in -15.5 dBFS · gain -4.5 dB · emolia-00167
(normally alert, steady, formal, newsreading)Um das Programm Zotero herunterzuladen, muss dieses auf der offiziellen Zotero-Seite heruntergeladen werden.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 6.5s, DE.
DE_1UW85J2KzTu_W000002 · in -16.6 dBFS · gain -3.4 dB · emolia-00167
(normally alert, steady, formal, didactic)Der Browser Connector aktiviert sich von selber. Sollte dies nicht der Fall sein, muss dies in den Einstellungen des Browsers manuell getan werden.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, didactic; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.5s, DE.
DE_1UW85J2KzTu_W000003 · in -19.6 dBFS · gain -0.4 dB · emolia-00167
(concentration· normally alert, steady, didactic, formal)Nach der Installation des Programms sollte in Word automatisch ein Menüpunkt mit dem Namen zu Tero auftauchen. Sollte dies nicht der Fall sein, muss manuell das Word Add-in nachinstalliert werden.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: didactic, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 12.1s, DE.
DE_1UW85J2KzTu_W000004 · in -18.1 dBFS · gain -1.9 dB · emolia-00167
This chain comes from the proxy rule: the same two-sided test as above, but because Contentment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Contentment clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.25.
At the same time Doubt goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.74 (higher than 74 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.02, then -0.05, then +0.05 — not a clean run: step 3 moves back the other way by 0.05 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.65 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.69 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.65, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 129 s · en · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.679 before conversion and 0.889 after — it rose by 0.209. Neighbour-to-neighbour the worst pair went 0.780 → 0.893. (The earlier render, with segment 1 left raw, scores 0.685 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.253 in the original and +0.162 after conversion — 64 % of the delta retained. On the other named axis, Doubt, -0.265 became -0.281.
Quality. Mean predicted overall quality across the segments went 2.90 → 3.02 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.679 → 0.889+0.209identity cos neighbours 0.780 → 0.893d_b rescored +0.253 → +0.162d_a rescored -0.265 → -0.281d_a mined -0.262d_b mined 0.252min_cos_consec (site) 0.6933min_cos_anchor (site) 0.6475dataset podcastlang enspeaker 951551total 127.8schain gain +4.7 dBseam step 1.1 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, quiet background, some disfluency, average clarity, moderate pitch range, light breath
(doubt, contemplation, sadness · normal-paced, subdued, neutral tension, casual)Yeah, I suppose just from the back of what Rose is saying there, I was just thinking about (ahem) um, yeah, the sort of wraparound is a horrible word. I don't really sound really weird. And so wraparound is what a term we use to mean things like (surprised gasp) um after show packs or like you know, workshops around a show as opposed to just the the piece itself.
full caption & clip details
A young adult feminine voice; delivery is subdued, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as doubt, contemplation, sadness; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 7.8/10; 29.7s, EN.
951551_00120093 · in -20.1 dBFS · gain +0.1 dB · podcast-03434
(hope enthusiasm optimism, concentration, contentment· normal-paced, subdued, slightly relaxed, casual)But it's something that at the head we've been thinking a lot about before the pandemic about how you introduce a story or work with a with a group of people before you make a piece of work and how you kind of yeah, start that process before the show happens, (ahem) and then also what happens afterwards. But I think this is just kind of unlocked, lots of other ways of doing that. On like maybe it's about making, you know, recorded footage or audio pieces afterwards or activity packs that you can do, you know. (ahem) Um there's like lots of different opportunities there.
full caption & clip details
A young adult feminine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as hope enthusiasm optimism, concentration, contentment; style: casual, monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 4.8/10; 29.8s, EN.
951551_00123064 · in -20.0 dBFS · gain +0.0 dB · podcast-03446
(jealousy and envy, hope enthusiasm optimism, contentment ·brisk, normally alert, neutral tension, casual)I've also just learned like really silly skills like captioning and (low mumble) uh editing videos (breathy giggle) and (ahem) um and other little things like that and and yeah, stuff (surprised gasp) that and even like exercises on Zoom and things like that. And like I wonder how I wonder if if when we go back to when the world opens up a bit more and things a bit normal, how we will will we still do some online options for people? Because actually maybe people they find it easier to engage when they don't have to travel and they
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as jealousy and envy, hope enthusiasm optimism, contentment; style: casual, conversational; good recording, quiet background; genuineness 4.1/6; vocal-burst blend 6.5/10; 27.1s, EN.
951551_00126040 · in -18.3 dBFS · gain -1.7 dB · podcast-06198
(contemplation, relief, contentment ·normal-paced, very low-energy, slightly relaxed, casual)don't have to leave the house. You know, there might be something about a blended, a blended future online and offline. Does feel like we're not just gonna go back. It's like a lot of these things will sit stay with us (surprised gasp) uh for a long time. (ahem) Um and that'll be that'll be (low mumble) (low mumble) really good.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contemplation, relief, contentment; style: casual, ASMR; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 5.7/10; 21.6s, EN.
951551_00128750 · in -20.8 dBFS · gain +0.8 dB · podcast-03442
(contentment, thankfulness gratitude, relief · normal-paced, normally alert, neutral tension, casual)So good practice. Um the (ahem) thing I saw the most recent thing I've seen, I'll just tell you about that because came to mind (ahem) uh straight away when you were asking that question. (ahem) Um was a show called The Origin of Carmen Power.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as contentment, thankfulness gratitude, relief; style: casual, conversational; good recording, quiet background; genuineness 4.6/6; vocal-burst blend 7.4/10; 20.2s, EN.
951551_00144185 · in -21.7 dBFS · gain +1.7 dB · podcast-03465
This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Contemplation clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.30.
At the same time Bitterness goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.13, then -0.04, then +0.03 — not a clean run: step 3 moves back the other way by 0.04 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.88 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 73 s · da · eurospeech
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.902 before conversion and 0.887 after — it fell by 0.016. Neighbour-to-neighbour the worst pair went 0.881 → 0.899. (The earlier render, with segment 1 left raw, scores 0.697 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.297 in the original and +0.245 after conversion — 83 % of the delta retained, which is most of it. On the other named axis, Bitterness, -0.361 became -0.073.
Quality. Mean predicted overall quality across the segments went 2.99 → 3.36 (+0.37) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.902 → 0.887-0.016identity cos neighbours 0.881 → 0.899d_b rescored +0.297 → +0.245d_a rescored -0.361 → -0.073d_a mined -0.362d_b mined 0.297min_cos_consec (site) 0.8807min_cos_anchor (site) 0.9047dataset eurospeechlang daspeaker denmark_20141M057_2015-02-total 72.0schain gain +2.5 dBseam step 0.3 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, slightly dark, slightly rough, balanced body, quiet background, measured, fairly steady, somewhat unclear
(normally alert, slightly relaxed, some disfluency, monologue)i vejen for virksomhederne, som det også er blevet sagt i forbindelse med nogle af høringssvarene, nemlig at man ligefrem skulle gøre det dårligere (low mumble) ved den måde, vi gør det på. Så det er ikke (low mumble) intentionen her.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 2.9/10; 12.3s, DA.
denmark_20141M057_2015-02-19_1000_6302000_6314255 · in -26.0 dBFS · gain +6.0 dB · eurospeech-00235
(awe, intoxication altered states of consciousness, concentration·very low-energy, relaxed, frequent disfluency, monologue)Omvendt skal det selvfølgelig heller ikke være sådan, at (low mumble) fordi vi har en branche, som lever godt af det, og som er dygtig – jeg er helt enig i beskrivelsen af, at et godt revideret regnskab er et vigtigt værktøj for enhver virksomhed og investor og andet –
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, submissive, slightly guarded; reads as awe, intoxication altered states of consciousness, concentration; style: monologue, casual; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 2.3/10; 18.1s, DA.
denmark_20141M057_2015-02-19_1000_6314255_6332400 · in -26.9 dBFS · gain +6.8 dB · eurospeech-00235
(contemplation, doubt, shame· very low-energy, neutral tension, frequent disfluency, monologue)Og jeg håber, at et flertal i Folketinget er med på at tænke på den måde. Vi vil selvfølgelig besvare alle spørgsmål positivt og holde en (low mumble) høring. Det gælder også den anden del af det, hvor jeg kan forstå at der er en vis skepsis, nemlig i forhold til CSR
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, doubt, shame; style: monologue, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 3.3/10; 17.6s, DA.
denmark_20141M057_2015-02-19_1000_6351744_6369376 · in -25.3 dBFS · gain +5.3 dB · eurospeech-00235
(triumph, pride, contemplation · very low-energy, relaxed, frequent disfluency, monologue)og det at implementere det (low mumble) under danske forhold. Jeg vil sige, at over 90 pct. af de virksomheder, som vi taler om her, har i forvejen CSR. Så det er jo noget, som danske virksomheder i allerhøjeste grad lever op til.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, slightly guarded; reads as triumph, pride, contemplation; style: monologue, whispered; below-average recording, quiet background; genuineness 3.7/6; vocal-burst blend 4.8/10; 14.3s, DA.
denmark_20141M057_2015-02-19_1000_6369376_6383696 · in -26.6 dBFS · gain +6.6 dB · eurospeech-00235
(contemplation · very low-energy, slightly relaxed, frequent disfluency, monologue)Og jeg synes egentlig, det er et ganske godt eksempel på et sted, hvor EU giver masser af mening og nytte, nemlig at vi sådan langsomt, men sikkert får
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation; style: monologue, whispered; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.1/10; 10.4s, DA.
denmark_20141M057_2015-02-19_1000_6383696_6394111 · in -25.2 dBFS · gain +5.2 dB · eurospeech-00235
This chain comes from the proxy rule: the same two-sided test as above, but because Contentment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Contentment clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.34.
At the same time Affection goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.10, then -0.06, then +0.10 — not a clean run: step 3 moves back the other way by 0.06 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.09 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.20 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.09, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 78 s · en · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.122 before conversion and 0.476 after — it rose by 0.354. Neighbour-to-neighbour the worst pair went 0.244 → 0.681. (The earlier render, with segment 1 left raw, scores 0.376 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.340 in the original and +0.411 after conversion — 121 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Affection, -0.208 became +0.589.
Quality. Mean predicted overall quality across the segments went 2.85 → 3.16 (+0.31) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.122 → 0.476+0.354identity cos neighbours 0.244 → 0.681d_b rescored +0.340 → +0.411d_a rescored -0.208 → +0.589d_a mined -0.358d_b mined 0.342min_cos_consec (site) 0.2029min_cos_anchor (site) 0.0888dataset podcastlang enspeaker 430744total 76.4schain gain +4.5 dBseam step 1.6 dBcrossfades 100/100/150/100 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · quiet background, moderately variable
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, neutral stance, neutral openness; reads as affection, elation, pleasure ecstasy; style: casual, conversational; good recording, quiet background; genuineness 3.9/6; vocal-burst blend 2.6/10; 4.4s, EN.
430744_00032952 · in -23.6 dBFS · gain +3.6 dB · podcast-01181
(interest, pleasure ecstasy, elation · normal-paced, normally alert, neutral tension, casual)I really had a lot of ideas for what this episode really should be called. So, Marcy, why don't you go first? Like what would you have called the actual episode in our own sassy speak?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as interest, pleasure ecstasy, elation; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 7.8/10; 11.4s, EN.
430744_00038338 · in -26.0 dBFS · gain +6.0 dB · podcast-05954
(pleasure ecstasy, elation, interest ·brisk, energised, neutral tension, casual)So I actually really like what is America because I thought of Sasha Barriconin's show. (ahem) Um, and I feel like everyone's just kind of like shrugging their shoulders and saying, like, what is this? Right. (ahem) Um, but after watching the first episode, (low mumble) um, and I'll go into this a little bit later, I think I would have called this, oh shit, this is the worst cult ever.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as pleasure ecstasy, elation, interest; style: casual, conversational; good recording, quiet background; genuineness 4.3/6; vocal-burst blend 7.3/10; 20.1s, EN.
430744_00039512 · in -25.4 dBFS · gain +5.4 dB · podcast-04737
(interest, hope enthusiasm optimism, elation ·normal-paced, energised, neutral tension, casual)Seriously, (ahem) I think if there's two things that stood out from this episode, it's who who works Purge Night and there are better cults to join. So let's break down this episode, John.
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as interest, hope enthusiasm optimism, elation; style: casual, conversational; below-average recording, quiet background; genuineness 4.3/6; vocal-burst blend 2.1/10; 13.6s, EN.
430744_00042576 · in -27.1 dBFS · gain +7.1 dB · podcast-04744
(contentment, elation, interest · normal-paced, very low-energy, neutral tension, conversational)Let's really break it down. So (low mumble) um what we really see, you know, is that it is (low mumble) um was well was filmed in New Orleans. (low mumble) Um, and right away when we are introduced to the opening credits, we're introduced to a character named Miguel. Marcy, do you want to tell us a little bit about him?
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as contentment, elation, interest; style: conversational, casual; below-average recording, quiet background; genuineness 3.6/6; vocal-burst blend 5.3/10; 27.4s, EN.
430744_00043984 · in -26.1 dBFS · gain +6.1 dB · podcast-05951
This chain comes from the proxy rule: the same two-sided test as above, but because Pride is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Pride clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.25.
At the same time Interest goes the other way, from 0.92 (higher than 92 % of clips in this corpus) to 0.62 (higher than 62 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.05, then -0.03, then +0.08, then +0.15 — not a clean run: step 2 moves back the other way by 0.03 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.91 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 72 s · en · eurospeech
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.846 before conversion and 0.870 after — it rose by 0.025. Neighbour-to-neighbour the worst pair went 0.846 → 0.870. (The earlier render, with segment 1 left raw, scores 0.740 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.251 in the original and +0.214 after conversion — 85 % of the delta retained, which is most of it. On the other named axis, Interest, -0.300 became -0.315.
Quality. Mean predicted overall quality across the segments went 3.08 → 3.28 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.846 → 0.870+0.025identity cos neighbours 0.846 → 0.870d_b rescored +0.251 → +0.214d_a rescored -0.300 → -0.315d_a mined -0.301d_b mined 0.251min_cos_consec (site) 0.9146min_cos_anchor (site) 0.8970dataset eurospeechlang enspeaker uk_uk_26_10072023total 71.0schain gain +2.6 dBseam step 1.0 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, normally alert, slightly relaxed, moderate pitch range
(interest, concentration, hope enthusiasm optimism · normal-paced, fairly steady, some disfluency, conversational)(low mumble) We have already examined in some detail the challenges (ahem) in the education system. We really need to get things moving and modernised in Northern Ireland.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, concentration, hope enthusiasm optimism; style: conversational, casual; good recording, no background noise; genuineness 3.8/6; vocal-burst blend 3.3/10; 10.4s, EN.
uk_uk_26_10072023_21727152_21737520 · in -24.8 dBFS · gain +4.8 dB · eurospeech-00981
(concentration · normal-paced, fairly steady, some disfluency, formal)In my view, that should come from a partnership between the Westminster Government and (low mumble) Stormont. We should all be working together to focus on the big issues, (low mumble) because people’s needs depend on it. (ahem) That is why we must urgently get over the hurdles
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration; style: formal, monologue; good recording, quiet background; genuineness 2.1/6; vocal-burst blend 1.8/10; 17.6s, EN.
uk_uk_26_10072023_21737520_21755136 · in -25.4 dBFS · gain +5.4 dB · eurospeech-00981
(emotional numbness, contentment·measured, steady, little disfluency, newsreading)So far, that has included Engage, Healthy Happy Minds, the school holiday food grant scheme and many more. However, significantly, a range of early years programmes will continue—thank goodness.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, little disfluency, moderate pitch range, minimal breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness, contentment; style: newsreading, whispered; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.6/10; 16.2s, EN.
uk_uk_26_10072023_21785536_21801696 · in -25.2 dBFS · gain +5.2 dB · eurospeech-00981
(concentration· measured, fairly steady, little disfluency, dramatic)That is after the Department produced an analysis of the impact that ending them would have on people’s lives. In the words of Dr Browne: “In considering the scale and cumulative impact of the proposed cuts,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, little disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration; style: dramatic, storytelling; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.2/10; 15.1s, EN.
uk_uk_26_10072023_21801696_21816784 · in -25.5 dBFS · gain +5.5 dB · eurospeech-00981
(pride, concentration, triumph· measured, fairly steady, almost no disfluency, newsreading)which represent a major change to long standing Ministerial programmes and policies, I am of the view that such a decision should be taken by a Minister, not a Permanent Secretary.”
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pride, concentration, triumph; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.8/10; 12.5s, EN.
uk_uk_26_10072023_21816784_21829311 · in -25.7 dBFS · gain +5.7 dB · eurospeech-00981
This chain comes from the proxy rule: the same two-sided test as above, but because Fatigue Exhaustion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Fatigue Exhaustion around average — 0.55, higher than 55 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.29.
At the same time Pleasure Ecstasy goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.63 (higher than 63 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.12, then +0.12, then -0.10, then +0.15 — not a clean run: step 3 moves back the other way by 0.10 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.04 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.08 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.04, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 60 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.056 before conversion and 0.506 after — it rose by 0.450. Neighbour-to-neighbour the worst pair went 0.056 → 0.549. (The earlier render, with segment 1 left raw, scores 0.350 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
The emotional move did not survive. Re-scored end to end, Fatigue Exhaustion moved +0.292 in the original and -0.100 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Pleasure Ecstasy, -0.329 became -0.285.
Quality. Mean predicted overall quality across the segments went 2.75 → 3.01 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.056 → 0.506+0.450identity cos neighbours 0.056 → 0.549d_b rescored +0.292 → -0.100d_a rescored -0.329 → -0.285d_a mined -0.329d_b mined 0.292min_cos_consec (site) 0.0786min_cos_anchor (site) 0.0409dataset emolialang enspeaker EN_Z-fIkCDoTXgtotal 58.7schain gain +4.4 dBseam step 1.5 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-bright
(pleasure ecstasy, hope enthusiasm optimism, embarrassment · normal-paced, normally alert, relaxed, casual)Yeah. So I'm the, I'm Nick. I'm the creative video architect (ahem) at InnoDeuce. Yeah. So I make all the videos that aren't live streams, except for the ones that sometimes James edits, but yeah. And Kat. And Kat.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, slightly guarded; reads as pleasure ecstasy, hope enthusiasm optimism, embarrassment; style: casual, monologue; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 4.5/10; 15.0s, EN.
EN_Z-fIkCDoTXg_W000014 · in -21.9 dBFS · gain +1.9 dB · emolia-02517
(contentment, intoxication altered states of consciousness, hope enthusiasm optimism · normal-paced, subdued, slightly relaxed, casual)Yeah, Nick and Kat and James. All right. Kat was just in here earlier this morning, finishing up, putting the finishing touches on, on one of the videos. So in what video will that be? We gotta stay tuned to our social media to find out what we'll be launching soon and what's done and all that other stuff. So follow us on (low mumble) Instagram.
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contentment, intoxication altered states of consciousness, hope enthusiasm optimism; style: casual, conversational; below-average recording, quiet background; genuineness 4.0/6; vocal-burst blend 6.7/10; 22.5s, EN.
EN_Z-fIkCDoTXg_W000015 · in -18.9 dBFS · gain -1.1 dB · emolia-02517
(hope enthusiasm optimism, pride· normal-paced, normally alert, slightly relaxed, casual)at introduce media. (low mumble) Uhm, obviously also introduce multimedia on both Facebook and YouTube and LinkedIn. So that's where we are. If you can't find us, you're not looking.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, pride; style: casual, conversational; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 2.8/10; 11.5s, EN.
EN_Z-fIkCDoTXg_W000016 · in -18.4 dBFS · gain -1.6 dB · emolia-02517
(normal-paced, normally alert, slightly relaxed, casual)That's all I'm gonna say about that. But anyway, on to the next at-
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: casual, conversational; average recording, some background noise; genuineness 4.8/6; vocal-burst blend 4.2/10; 3.7s, EN.
EN_Z-fIkCDoTXg_W000017 · in -17.9 dBFS · gain -2.1 dB · emolia-02517
(measured, normally alert, slightly relaxed, monologue)(low mumble) Uhm, episode, on to the next thing, which is what, what we talk about, what we're working on. So cue it.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; good recording, quiet background; genuineness 3.0/6; vocal-burst blend 1.8/10; 6.9s, EN.
EN_Z-fIkCDoTXg_W000018 · in -19.6 dBFS · gain -0.4 dB · emolia-02517
This chain comes from the proxy rule: the same two-sided test as above, but because Pain is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Pain around average — 0.54, higher than 54 % of clips in this corpus — and ends with it strongly present at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.28.
At the same time Doubt goes the other way, from 0.83 (higher than 83 % of clips in this corpus) to 0.55 (higher than 55 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.00, then +0.00, then +0.17, then +0.10 — a plateau around step 1, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.72 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.82 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.72, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 36 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.701 before conversion and 0.819 after — it rose by 0.118. Neighbour-to-neighbour the worst pair went 0.724 → 0.807. (The earlier render, with segment 1 left raw, scores 0.754 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.276 in the original and +0.292 after conversion — 106 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Doubt, -0.272 became -0.515.
Quality. Mean predicted overall quality across the segments went 2.91 → 3.19 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.701 → 0.819+0.118identity cos neighbours 0.724 → 0.807d_b rescored +0.276 → +0.292d_a rescored -0.272 → -0.515d_a mined -0.272d_b mined 0.276min_cos_consec (site) 0.8178min_cos_anchor (site) 0.7179dataset emolialang zhspeaker ZH_B00000_S01555total 34.4schain gain -0.8 dBseam step 1.2 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, slightly dark, fairly smooth, balanced body, no background noise, measured, slightly relaxed, no disfluency
This chain comes from the proxy rule: the same two-sided test as above, but because Jealousy and Envy is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Jealousy and Envy clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.32.
At the same time Doubt goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.04, then -0.05, then +0.12 — not a clean run: step 3 moves back the other way by 0.05 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 125 s · en · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.912 before conversion and 0.880 after — it fell by 0.033. Neighbour-to-neighbour the worst pair went 0.929 → 0.898. (The earlier render, with segment 1 left raw, scores 0.721 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Jealousy and Envy moved +0.315 in the original and +0.183 after conversion — 58 % of the delta retained. On the other named axis, Doubt, -0.440 became -0.627.
Quality. Mean predicted overall quality across the segments went 3.14 → 3.23 (+0.09) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.912 → 0.880-0.033identity cos neighbours 0.929 → 0.898d_b rescored +0.315 → +0.183d_a rescored -0.440 → -0.627d_a mined -0.396d_b mined 0.322min_cos_consec (site) 0.9474min_cos_anchor (site) 0.9385dataset podcastlang enspeaker 501941total 124.1schain gain +4.6 dBseam step 0.6 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · fairly smooth, slightly thin, average recording, quiet background, average clarity
(doubt, concentration, thankfulness gratitude · normal-paced, subdued, slightly relaxed, casual)From my perspective, I think that there is I mean, statistically, I'll talk statistically. The numbers say that female founders raise in women's health, they raise about two to three percent of capital. It has been changing. So there is clearly an opportunity there to improve it, especially considering that according this is all statistics, so it's all published, it's not me saying it.
full caption & clip details
A young adult feminine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as doubt, concentration, thankfulness gratitude; style: casual, monologue; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 2.4/10; 27.4s, EN.
501941_00125416 · in -17.7 dBFS · gain -2.3 dB · podcast-05128
(hope enthusiasm optimism, interest, pride·brisk, normally alert, neutral tension, casual)(ahem) Um statistically, they have better (low mumble) um rates of success. That like we were talking about product market fit, we tend to build great communities that follow our products. So I think there's a huge opportunity to put more money here if we want real returns, not only from the side of economic returns, but also social re returns.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism, interest, pride; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 9.9/10; 23.6s, EN.
501941_00128159 · in -17.6 dBFS · gain -2.5 dB · podcast-05133
(doubt, hope enthusiasm optimism, elation·normal-paced, very low-energy, slightly relaxed, casual)And I mean, it was last year where the World Economic Forum mentioned that for every $1 that we invest in women's health, we get $3 back into the economy. So I mean. Statistically, it just makes sense. Me personally, I think the biggest challenge is that what I've seen also from my cohort of uh (low mumble) Femtech founders is that many of these companies don't necessarily fit the checklist that (low mumble) uh
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as doubt, hope enthusiasm optimism, elation; style: casual, monologue; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 2.4/10; 30.0s, EN.
501941_00130511 · in -18.1 dBFS · gain -1.9 dB · podcast-05128
(interest, contemplation, concentration· normal-paced, subdued, slightly relaxed, casual)traditional venture capital looks like. Uh, (ahem) because sometimes our clinical developments that take time. Sometimes there are products that are in between medical but come (ahem) uh consumer focus. So is it one or the other? Or sometimes the development times, yeah, maybe too long or maybe too short. So all sorts of combinations, I would say.
full caption & clip details
A young adult feminine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as interest, contemplation, concentration; style: casual, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 4.1/10; 23.8s, EN.
501941_00133504 · in -17.6 dBFS · gain -2.4 dB · podcast-05131
(jealousy and envy, pride, affection· normal-paced, normally alert, neutral tension, casual)Me personally, what has worked very well is that we have built a community of supporters. Initially, a lot of women, because women understand the need and they're like, Yeah, I've been through that. And every time you share the story, there's always someone that will tell you, Oh, yeah, by the way, this this is my and they get interested.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as jealousy and envy, pride, affection; style: casual, monologue; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 8.3/10; 19.9s, EN.
501941_00135888 · in -17.3 dBFS · gain -2.7 dB · podcast-05143
This chain comes from the proxy rule: the same two-sided test as above, but because Sourness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Sourness clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.30.
At the same time Hope Enthusiasm Optimism goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.10, then -0.10, then +0.17, then +0.13 — not a clean run: step 2 moves back the other way by 0.10 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.18 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.19 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.18, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 60 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.284 before conversion and 0.650 after — it rose by 0.366. Neighbour-to-neighbour the worst pair went 0.291 → 0.607. (The earlier render, with segment 1 left raw, scores 0.539 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sourness moved +0.304 in the original and +0.362 after conversion — 119 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Hope Enthusiasm Optimism, -0.289 became -0.302.
Quality. Mean predicted overall quality across the segments went 2.87 → 2.99 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.284 → 0.650+0.366identity cos neighbours 0.291 → 0.607d_b rescored +0.304 → +0.362d_a rescored -0.289 → -0.302d_a mined -0.289d_b mined 0.304min_cos_consec (site) 0.1938min_cos_anchor (site) 0.1779dataset emolialang enspeaker EN_UAZH21Q84fytotal 58.8schain gain +6.2 dBseam step 1.9 dBcrossfades 100/150/150/100 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a middle-aged somewhat feminine voice · neutral-toned, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range, light breath
(hope enthusiasm optimism, contemplation, elation · normal-paced, little disfluency, clear, monologue)us imagine that future that he wants for us. And that is what I hope we can all do for ourselves.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, contemplation, elation; style: monologue, authoritative; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.0/10; 7.8s, EN.
EN_UAZH21Q84fy_W000166 · in -20.9 dBFS · gain +0.9 dB · emolia-01764
(contentment, hope enthusiasm optimism, contemplation · normal-paced, some disfluency, average clarity, casual)So instead of just knowing in our gut that we want this situation to go away, or we want to put an end to it, or we wish it was done, or we wish we could be free from it, let's get really specific. What exactly do I want this future to look like? That's not easy. It's not as easy as it might sound, but we can use Dr. King's I have a dream speech as a great example.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contentment, hope enthusiasm optimism, contemplation; style: casual, conversational; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 2.9/10; 22.6s, EN.
EN_UAZH21Q84fy_W000167 · in -23.8 dBFS · gain +3.8 dB · emolia-01764
(awe, contentment, thankfulness gratitude·measured, some disfluency, average clarity, monologue)And your book offers wonderful prompts in that chapter for sort of writing about, (low mumble) you know, to use your imagination.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as awe, contentment, thankfulness gratitude; style: monologue, conversational; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 0.0/10; 7.8s, EN.
EN_UAZH21Q84fy_W000168 · in -20.3 dBFS · gain +0.3 dB · emolia-01764
(concentration, interest, contentment ·normal-paced, some disfluency, average clarity, conversational)Yes, yes. One thing on the imagination distinction, I talk about it as using your imagination because if your brain, if your thinking brain could solve this situation and you've been working for years and years and years to do it, or even for months or even weeks, your brain
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as concentration, interest, contentment; style: conversational, casual; good recording, quiet background; genuineness 2.6/6; vocal-burst blend 2.9/10; 17.7s, EN.
EN_UAZH21Q84fy_W000170 · in -22.3 dBFS · gain +2.3 dB · emolia-01764
(normal-paced, some disfluency, average clarity, casual)Probably would have figured it out right by now, right? If you've learned all the best.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, conversational; average recording, no background noise; genuineness 3.6/6; vocal-burst blend 4.0/10; 3.6s, EN.
EN_UAZH21Q84fy_W000171 · in -20.5 dBFS · gain +0.5 dB · emolia-01764