Manifest tier. merged, rule B1 UNION VN1, T=0.2, step cap 0.25. Population 0 chains (0 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 0.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words
are a real non-speech sound, printed where it happens.
The full generated caption for any clip is under “full caption & clip
details”. Its perceived-gender and background-noise clauses were re-rendered from
the numeric buckets, because the versions stored in the corpus index had those two
ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
The tags and captions describe the ORIGINAL clips — they are the
corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the
converted audio are the ones printed in each card's own paragraph and chips.
The chain starts with Concentration clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.28.
It takes 3 clips to get there. Clip to clip the moves are +0.06, then +0.22 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 48 s · dutch · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.958 before conversion and 0.941 after — it fell by 0.017. Neighbour-to-neighbour the worst pair went 0.955 → 0.941. (The earlier render, with segment 1 left raw, scores 0.764 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.284 in the original and +0.405 after conversion — 143 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Awe, -0.027 became -0.474.
Quality. Mean predicted overall quality across the segments went 3.17 → 3.40 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.958 → 0.941-0.017identity cos neighbours 0.955 → 0.941d_b rescored +0.284 → +0.405d_a rescored -0.027 → -0.474d_a mined -0.027d_b mined 0.285min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang dutchspeaker 1724total 46.9schain gain +1.3 dBseam step 0.9 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady
(awe, longing, sexual lust · narration, monologue)zijn ogen die half onder de wenkbrauwen gedoken waren droegen het kenmerk ener warme en eenzame ziel alhoewel hij in doorluchtigheid voor de andere ridders niet wijken moest bleef hij nochtans achteruit en liet zijn minderen voorgaan
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, longing, sexual lust; style: narration, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 13.1s, DUTCH.
1724_1878_002627 · in -29.6 dBFS · gain +9.6 dB · mls-00100
(doubt, anger, thankfulness gratitude·storytelling, narration)meermalen had men de plaats geruimd om hem door te laten doch hij gaf geen acht op deze beleefdheid en draaide gedurig het hoofd naar de vrouwenstoet om zeker was er onder haar een lief beeld dat zijn hart aan een leiband hield en tot zich trok want men zag op zijn gelaat een zuivere glimlach verschijnen zodra hij het hoofd tot de vrouwen gewend had
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, anger, thankfulness gratitude; style: storytelling, narration; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 0.6/10; 19.5s, DUTCH.
1724_1878_002685 · in -30.1 dBFS · gain +10.1 dB · mls-00100
(concentration, awe, malevolence malice· narration, storytelling)bij het eerste gezicht zou men deze adolf voor een zoon van robrecht van bethune kunnen aanzien hebben want uitgenomen de ouderdom die zeer verschillend was geleek hij verwonderlijk aan robrecht dezelfde gestalte dezelfde houding dezelfde trekken in het gelaat
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, awe, malevolence malice; style: narration, storytelling; good recording, quiet background; genuineness 0.2/6; vocal-burst blend 0.0/10; 14.7s, DUTCH.
1724_1878_002236 · in -29.5 dBFS · gain +9.5 dB · mls-00100
The chain starts with style: asmr (S_ASMR) low — 0.09, lower than 91 % of clips in this corpus — and ends with it below average at 0.33, lower than 67 % of clips in this corpus. That is a total rise of 0.25.
It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.93 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.93 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 20 s · en · emolia
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.825 before conversion and 0.791 after — it fell by 0.033. Neighbour-to-neighbour the worst pair went 0.825 → 0.791. (The earlier render, with segment 1 left raw, scores 0.683 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, S_ASMR — style: asmr moved +0.246 in the original and +0.222 after conversion — 90 % of the delta retained, which is essentially all of it.
Quality. Mean predicted overall quality across the segments went 2.87 → 3.04 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.825 → 0.791-0.033identity cos neighbours 0.825 → 0.791d_b rescored +0.246 → +0.222d_a rescored +0.246 → +0.222d_a mined 0.247d_b mined 0.247min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00055_S00031total 19.4schain gain +3.0 dBseam step 1.5 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normally alert, slightly relaxed, some disfluency
(concentration, interest · normal-paced, fairly steady, average clarity, casual)First of all, let's back up again. What is a clause? (ahem) A clause is just a language chunk that has a subject and a verb. That's what a clause does.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, interest; style: casual, authoritative; good recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.3/10; 9.2s, EN.
EN_B00055_S00031_W000003 · in -17.9 dBFS · gain -2.1 dB · emolia-01297
(emotional numbness, contemplation, concentration ·measured, steady, clear, didactic)All sentences are clauses but not all clauses are sentences. I'll write that down. So all sentences are clauses but not all clauses are sentences.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, contemplation, concentration; style: didactic, authoritative; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 0.5/10; 10.4s, EN.
EN_B00055_S00031_W000004 · in -17.5 dBFS · gain -2.5 dB · emolia-01297
The chain starts with style: monologue (S_MONO) above average — 0.65, higher than 65 % of clips in this corpus — and works its way down to below average at 0.37, lower than 63 % of clips in this corpus. That is a total fall of 0.28.
It takes 3 clips to get there. Clip to clip the moves are -0.21, then -0.08 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.85 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.85 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 39 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.795 before conversion and 0.711 after — it fell by 0.085. Neighbour-to-neighbour the worst pair went 0.796 → 0.693. (The earlier render, with segment 1 left raw, scores 0.725 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, S_MONO — style: monologue moved -0.283 in the original and -0.229 after conversion — 81 % of the delta retained, which is most of it.
Quality. Mean predicted overall quality across the segments went 3.06 → 3.29 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.795 → 0.711-0.085identity cos neighbours 0.796 → 0.693d_b rescored -0.283 → -0.229d_a rescored -0.283 → -0.229d_a mined -0.284d_b mined -0.284min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00099_S00328total 37.8schain gain +3.2 dBseam step 1.3 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, fairly smooth, balanced body, quiet background
(concentration, contemplation, doubt · slow, very low-energy, relaxed, monologue)Sure, so, so first of all, the, the idea of, of a, of a Jew and becoming a Jew or a Jewish soul, remember that when, when a convert converts,
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, submissive, neutral openness; reads as concentration, contemplation, doubt; style: monologue, casual; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 0.2/10; 11.8s, EN.
EN_B00099_S00328_W000000 · in -23.5 dBFS · gain +3.5 dB · emolia-02143
(interest, shame, disappointment·brisk, energised, neutral tension, casual)To correct a specific thing that this righteous, this sadist, this righteous individual thought that if you knew these specific details, it would enhance you in this life. It was all about this life. It wasn't just, oh, I went to a conference and I got hypnotized and decided, you know, to find out who I was. I was Sri Lankan actually. But, (ahem) uh, but, (low mumble) um, the idea is,
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as interest, shame, disappointment; style: casual, dramatic; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 6.5/10; 19.5s, EN.
EN_B00099_S00328_W000001 · in -24.1 dBFS · gain +4.1 dB · emolia-02143
(emotional numbness·normal-paced, normally alert, slightly relaxed, conversational)My life scenarios with the family that was born into has nothing to do with the genetics that a person has.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as emotional numbness; style: conversational, casual; good recording, quiet background; genuineness 3.5/6; vocal-burst blend 3.1/10; 7.0s, EN.
EN_B00099_S00328_W000002 · in -23.6 dBFS · gain +3.6 dB · emolia-02143
The chain starts with Jealousy and Envy around average — 0.55, higher than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.38.
It takes 5 clips to get there. Clip to clip the moves are -0.10, then +0.25, then +0.05, then +0.18 — not a clean run: step 1 moves back the other way by 0.10 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 64 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.819 before conversion and 0.847 after — it rose by 0.028. Neighbour-to-neighbour the worst pair went 0.834 → 0.846. (The earlier render, with segment 1 left raw, scores 0.795 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Jealousy and Envy moved +0.375 in the original and +0.169 after conversion — 45 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Contempt, -0.225 became -0.001.
Quality. Mean predicted overall quality across the segments went 3.05 → 3.21 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.819 → 0.847+0.028identity cos neighbours 0.834 → 0.846d_b rescored +0.375 → +0.169d_a rescored -0.225 → -0.001d_a mined -0.224d_b mined 0.376min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00077_S08054total 62.2schain gain +3.0 dBseam step 1.9 dBcrossfades 150/150/100/100 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady
(contempt, impatience and irritability, disgust · clear, monologue, didactic)这是他们视野当中唯一的东西都不说是唯一重要的了。我觉得这好像就是他们生活当中唯一的一件事情,一个重大的事项。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contempt, impatience and irritability, disgust; style: monologue, didactic; average recording, no background noise; genuineness 2.7/6; vocal-burst blend 5.7/10; 10.2s, ZH.
ZH_B00077_S08054_W000047 · in -20.2 dBFS · gain +0.2 dB · emolia-04046
The chain starts with Concentration clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.32.
It takes 5 clips to get there. Clip to clip the moves are +0.16, then +0.22, then -0.25, then +0.19 — not a clean run: step 3 moves back the other way by 0.25 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.02 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.02 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 47 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.157 before conversion and 0.455 after — it rose by 0.298. Neighbour-to-neighbour the worst pair went 0.215 → 0.630. (The earlier render, with segment 1 left raw, scores 0.384 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.322 in the original and +0.351 after conversion — 109 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contempt, -0.380 became -0.249.
Quality. Mean predicted overall quality across the segments went 2.67 → 3.01 (+0.33) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.157 → 0.455+0.298identity cos neighbours 0.215 → 0.630d_b rescored +0.322 → +0.351d_a rescored -0.380 → -0.249d_a mined -0.380d_b mined 0.322min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_WczYyO4WQxUtotal 45.3schain gain +3.0 dBseam step 1.2 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · fairly smooth
(contempt, impatience and irritability · brisk, energised, neutral tension, authoritative)Will the Premier listen to the scientists and not the land speculators and cancel Highway 413?
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, fairly guarded; reads as contempt, impatience and irritability; style: authoritative, formal; average recording, quiet background; genuineness 0.1/6; vocal-burst blend 0.4/10; 7.5s, EN.
EN_WczYyO4WQxU_W000248 · in -19.4 dBFS · gain -0.6 dB · emolia-02125
(embarrassment, thankfulness gratitude·normal-paced, normally alert, slightly relaxed, formal)Earlier to the member opposite, I thank the member for the question. As I indicated earlier
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as embarrassment, thankfulness gratitude; style: formal, playful; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.9/10; 4.2s, EN.
EN_WczYyO4WQxU_W000251 · in -19.2 dBFS · gain -0.8 dB · emolia-02125
(concentration, triumph, interest·brisk, normally alert, slightly relaxed, formal)(ahem) We are doing the important work of completing a complete, a fulsome environmental assessment process as we consider whether or not to proceed with Highway 413. We believe, unlike the Liberals, that it's important to collect all the evidence. The information that was published today is of course of great interest and will feed into the work that we're doing.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is slightly cool, slightly bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, minimal breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration, triumph, interest; style: formal, authoritative; average recording, quiet background; genuineness 0.8/6; vocal-burst blend 1.4/10; 17.6s, EN.
EN_WczYyO4WQxU_W000252 · in -21.2 dBFS · gain +1.2 dB · emolia-02125
(brisk, normally alert, slightly relaxed, formal)But we believe that the demographic growth in the greater golden horseshoe to come in the next few decades
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is slightly cool, slightly bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: formal, authoritative; average recording, quiet background; genuineness 0.3/6; vocal-burst blend 1.5/10; 6.3s, EN.
EN_WczYyO4WQxU_W000253 · in -20.1 dBFS · gain +0.1 dB · emolia-02125
(concentration· brisk, normally alert, slightly relaxed, formal)Warrants, our government taking the time to consider what the transportation needs are of the greater Golden Horseshoe, and that means continuing with the environmental assessment process for Highway 413.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: formal, authoritative; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.3/10; 10.5s, EN.
EN_WczYyO4WQxU_W000254 · in -22.7 dBFS · gain +2.7 dB · emolia-02125
The chain starts with Thankfulness Gratitude clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.23.
It takes 4 clips to get there. Clip to clip the moves are +0.03, then +0.05, then +0.15 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.02 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.07 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.02, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 52 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.016 before conversion and 0.556 after — it rose by 0.540. Neighbour-to-neighbour the worst pair went 0.065 → 0.581. (The earlier render, with segment 1 left raw, scores 0.439 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.246 in the original and +0.290 after conversion — 118 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Jealousy and Envy, -0.658 became -0.432.
Quality. Mean predicted overall quality across the segments went 2.95 → 3.11 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.016 → 0.556+0.540identity cos neighbours 0.065 → 0.581d_b rescored +0.246 → +0.290d_a rescored -0.658 → -0.432d_a mined -0.657d_b mined 0.233min_cos_consec (site) 0.0658min_cos_anchor (site) 0.0188dataset podcastlang enspeaker 204298total 51.5schain gain +4.8 dBseam step 1.7 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · average clarity
(jealousy and envy, astonishment surprise, confusion · normal-paced, very low-energy, relaxed, casual)I don't realize or don't think about how cool some of the experiences that we've had. Like how people, how is your summer? Oh yeah, it's fine. We actually did some really cool stuff this summer.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, slightly vulnerable; reads as jealousy and envy, astonishment surprise, confusion; style: casual, conversational; good recording, no background noise; genuineness 3.0/6; vocal-burst blend 2.6/10; 13.2s, EN.
204298_00141704 · in -28.1 dBFS · gain +8.1 dB · podcast-01084
(relief, embarrassment, contentment·slow, very low-energy, relaxed, whispered)Yeah, it was cool. It was a good summer. But anyways. And the weather was good. And business was decent. It
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as relief, embarrassment, contentment; style: whispered, monologue; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 1.8/10; 8.9s, EN.
204298_00143048 · in -29.7 dBFS · gain +9.7 dB · podcast-01081
(normal-paced, normally alert, slightly relaxed, casual)Yeah, the weather, the good weather across the country really helped with business because it was looking a little ugly at the beginning of the year, but it turned out okay.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 4.2/6; vocal-burst blend 5.3/10; 8.2s, EN.
204298_00144184 · in -27.2 dBFS · gain +7.2 dB · podcast-01070
(thankfulness gratitude, affection, hope enthusiasm optimism·measured, subdued, neutral tension, casual)What is your hope for the Canadian fishing podcast going (low mumble) forward? And (low mumble) let's let's end on that note.
full caption & clip details
An adult somewhat masculine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, neutral openness; reads as thankfulness gratitude, affection, hope enthusiasm optimism; style: casual, whispered; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.5/10; 21.8s, EN.
204298_00145576 · in -27.6 dBFS · gain +7.6 dB · podcast-01072
FULL — fullness of tone ↓identity −0.05emotion 28 % merged_emo_vn__T0.20__C0.25__INTERNAL · #7
The chain starts with fullness of tone (FULL) above average — 0.65, higher than 65 % of clips in this corpus — and works its way down to below average at 0.31, lower than 69 % of clips in this corpus. That is a total fall of 0.35.
It takes 4 clips to get there. Clip to clip the moves are -0.20, then -0.01, then -0.13 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 97 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.922 before conversion and 0.876 after — it fell by 0.046. Neighbour-to-neighbour the worst pair went 0.924 → 0.897. (The earlier render, with segment 1 left raw, scores 0.738 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, FULL — fullness of tone moved -0.348 in the original and -0.098 after conversion — 28 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 2.81 → 3.30 (+0.49) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.922 → 0.876-0.046identity cos neighbours 0.924 → 0.897d_b rescored -0.348 → -0.098d_a rescored -0.348 → -0.098d_a mined -0.346d_b mined -0.346min_cos_consec (site) 0.9444min_cos_anchor (site) 0.9290dataset podcastlang enspeaker 603760total 95.9schain gain +3.1 dBseam step 0.4 dBcrossfades 150/150/150 ms
Script — 4 chunks, 4 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-toned, slightly rough, relaxed, fairly steady, frequent disfluency, somewhat unclear
(contentment, elation, hope enthusiasm optimism · measured, very low-energy, fairly narrow pitch, casual)Yeah this this year especially it's been quite busy (low mumble) um and I'm fortunate enough that (ahem) uh my university they have a lot of (ahem) uh support for creative projects like recording funding. So it's really a really a uh (ahem) uh a unique situation because not not too many places have have that like regularly every year.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, slightly submissive, slightly guarded; reads as contentment, elation, hope enthusiasm optimism; style: casual, monologue; below-average recording, quiet background; genuineness 4.5/6; vocal-burst blend 4.6/10; 23.6s, EN.
603760_00305880 · in -16.8 dBFS · gain -3.2 dB · podcast-03863
(contentment, pride, hope enthusiasm optimism · measured, subdued, fairly narrow pitch, monologue)So (low mumble) um so this year I started a project with some of my f colleagues in the at the university we're doing recording three pieces. One of them is a (low mumble) um saxophone concerto uh (low mumble) by Kirk (low mumble) uh O'Reodon uh (low mumble) American composer that he wrote that piece actually for Dr. Eugene Rousseau's seventy sixth birthday celebration and so I premiered that piece and so finally I have a chance to to play that it's with chamber
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contentment, pride, hope enthusiasm optimism; style: monologue, casual; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 4.3/10; 28.7s, EN.
603760_00308240 · in -17.0 dBFS · gain -3.0 dB · podcast-03862
(pleasure ecstasy, elation, interest·slow, very low-energy, narrow pitch range, ASMR)chamber group concerto and then s (ahem) uh the other two are also chamber music like the creation of the world by Mio and the Kurtville, the three penny opera suite. And so that's a that's a I think exciting (low mumble) um
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, narrow pitch range, normal breath; affect is mildly negative, submissive, neutral openness; reads as pleasure ecstasy, elation, interest; style: ASMR, monologue; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 4.4/10; 15.1s, EN.
603760_00311106 · in -17.7 dBFS · gain -2.3 dB · podcast-03862
(doubt, intoxication altered states of consciousness, shame·measured, very low-energy, fairly narrow pitch, casual)I don't know if any in the university actually have faculty done recording like that together. So it I think a unique thing. And also I'm f still finishing my (ahem) uh sonata, my volume two sonata recording. I've I finished recording it just now editing. And then also started (low mumble) um another project is (low mumble) um just yesterday I was joking with somebody, I said, you know, talking about quart saxophone quartet.
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly negative, submissive, neutral openness; reads as doubt, intoxication altered states of consciousness, shame; style: casual, monologue; below-average recording, some background noise; genuineness 4.6/6; vocal-burst blend 8.9/10; 29.1s, EN.
603760_00312612 · in -18.0 dBFS · gain -2.0 dB · podcast-00354
The chain starts with Hope Enthusiasm Optimism clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.22.
It takes 2 clips to get there. Clip to clip the moves are +0.22 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.87 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.87 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 19 s · en · emolia
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.787 before conversion and 0.850 after — it rose by 0.063. Neighbour-to-neighbour the worst pair went 0.787 → 0.850. (The earlier render, with segment 1 left raw, scores 0.663 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.221 in the original and +0.184 after conversion — 83 % of the delta retained, which is most of it. On the other named axis, Contentment, -0.194 became -0.060.
Quality. Mean predicted overall quality across the segments went 2.75 → 3.07 (+0.32) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.787 → 0.850+0.063identity cos neighbours 0.787 → 0.850d_b rescored +0.221 → +0.184d_a rescored -0.194 → -0.060d_a mined -0.195d_b mined 0.221min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_785inmct89ktotal 18.2schain gain +1.3 dBseam step 3.1 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, slightly bright, fairly smooth, normal-paced, normally alert, slightly relaxed, moderate pitch range, light breath
(fairly steady, frequent disfluency, somewhat unclear, casual)Two of the tenets of this model, two of the connections are to use students' first names and to schedule one-on-one meetings with them. So you're taking care of two of those with these meetings. And (ahem) uhm,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.0/10; 11.7s, EN.
EN_785inmct89k_W000009 · in -20.9 dBFS · gain +0.9 dB · emolia-00988
(hope enthusiasm optimism·steady, some disfluency, clear, monologue)You could also encourage students to come in pairs in groups. This works really well for review sessions and project planning.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, slightly bright, fairly smooth, very thin; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as hope enthusiasm optimism; style: monologue, dramatic; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 1.3/10; 6.7s, EN.
EN_785inmct89k_W000010 · in -17.6 dBFS · gain -2.5 dB · emolia-00988
The chain starts with harmonicity (HARM) below average — 0.32, lower than 68 % of clips in this corpus — and works its way down to at the very bottom of the range at 0.04, lower than 96 % of clips in this corpus. That is a total fall of 0.28.
It takes 4 clips to get there. Clip to clip the moves are +0.02, then -0.11, then -0.20 — not a clean run: step 1 moves back the other way by 0.02 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.40 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.19 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.40, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 63 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.562 before conversion and 0.799 after — it rose by 0.237. Neighbour-to-neighbour the worst pair went 0.426 → 0.734. (The earlier render, with segment 1 left raw, scores 0.735 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, HARM — harmonicity moved -0.279 in the original and -0.158 after conversion — 57 % of the delta retained.
Quality. Mean predicted overall quality across the segments went 2.47 → 3.02 (+0.55) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.562 → 0.799+0.237identity cos neighbours 0.426 → 0.734d_b rescored -0.279 → -0.158d_a rescored -0.279 → -0.158d_a mined -0.283d_b mined -0.283min_cos_consec (site) 0.1910min_cos_anchor (site) 0.4012dataset podcastlang enspeaker 914894total 61.7schain gain +2.3 dBseam step 2.8 dBcrossfades 150/150/100 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · below-average recording, some background noise, energised, neutral tension, moderately variable, some disfluency, wide pitch range
(elation, hope enthusiasm optimism, embarrassment · brisk, average clarity, light breath, conversational)(ahem) (ahem) Machina (childlike giggle) international un saludo para Vanessa who is all the España
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, slightly guarded; reads as elation, hope enthusiasm optimism, embarrassment; style: conversational, casual; below-average recording, some background noise; genuineness 4.9/6; vocal-burst blend 7.0/10; 16.8s, EN.
914894_00277296 · in -21.4 dBFS · gain +1.4 dB · podcast-00079
(interest, sourness, disgust· brisk, somewhat unclear, light breath, casual)Hasta España. No, nada más para terminar antes de (ahem) Ivonne. Sur la commentary of Fernando Delírio, he mentioned that there's a moment in which we restring the consumer of the video. You don't know if we're going to this point with respect to create politics that restrict the consumer of one of the empresas more proliferative of the ultimate years. There are millions and millions of dinner that is generated in the empresas of the video.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as interest, sourness, disgust; style: casual, dramatic; below-average recording, some background noise; genuineness 4.2/6; vocal-burst blend 10.0/10; 26.5s, EN.
914894_00282552 · in -24.2 dBFS · gain +4.2 dB · podcast-01910
(teasing, pleasure ecstasy, elation·normal-paced, average clarity, normal breath, casual)De hecho, the two industries more in the pandemic are the videojuegs and the porno. Some are very necessary.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as teasing, pleasure ecstasy, elation; style: casual, conversational; below-average recording, some background noise; genuineness 6.0/6; vocal-burst blend 2.1/10; 8.8s, EN.
914894_00287472 · in -20.2 dBFS · gain +0.2 dB · podcast-00046
(affection, thankfulness gratitude, pleasure ecstasy ·brisk, somewhat unclear, normal breath, casual)for the money of all the problems that we have in the pandemic in the things that for example Spanish that were super limited,
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as affection, thankfulness gratitude, pleasure ecstasy; style: casual, playful; below-average recording, some background noise; genuineness 5.6/6; vocal-burst blend 10.0/10; 10.1s, EN.
914894_00291264 · in -25.9 dBFS · gain +6.0 dB · podcast-00048
The chain starts with Bitterness clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.25.
It takes 5 clips to get there. Clip to clip the moves are +0.22, then -0.04, then -0.01, then +0.08 — not a clean run: step 2 moves back the other way by 0.04 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 76 s · pt · eurospeech
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.213 before conversion and 0.694 after — it rose by 0.481. Neighbour-to-neighbour the worst pair went 0.233 → 0.792. (The earlier render, with segment 1 left raw, scores 0.664 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Bitterness moved +0.254 in the original and +0.353 after conversion — 139 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Thankfulness Gratitude, -0.023 became -0.894.
Quality. Mean predicted overall quality across the segments went 3.05 → 3.22 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.213 → 0.694+0.481identity cos neighbours 0.233 → 0.792d_b rescored +0.254 → +0.353d_a rescored -0.023 → -0.894d_a mined -0.023d_b mined 0.254min_cos_consec (site) —min_cos_anchor (site) —dataset eurospeechlang ptspeaker portugal_15_1_153total 74.6schain gain +3.9 dBseam step 1.0 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · slightly cool, neutral-bright, average recording, quiet background, moderately variable, almost no disfluency, wide pitch range
(thankfulness gratitude · brisk, normally alert, neutral tension, dramatic)entendemos que o grande desígnio desta Legislatura é melhorar os salários e as remunerações dos nossos trabalhadores. Aplausos do PS. O Sr. Presidente:
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, almost no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as thankfulness gratitude; style: dramatic, authoritative; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 3.3/10; 11.6s, PT.
portugal_15_1_153_2098544_2110129 · in -21.2 dBFS · gain +1.2 dB · eurospeech-02615
(anger, concentration, malevolence malice·measured, energised, slightly relaxed, cartoonish)O Sr. Rui Paulo Sousa (CH): — … um ataque encabeçado pelo seu responsável máximo, o Sr. Primeiro- Ministro, que sempre achou que as ordens têm demasiado poder. Mas o que isto quer dizer, na verdade, é que não se vergam com facilidade ao poder socialista.
full caption & clip details
A middle-aged masculine voice; delivery is energised, measured, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, rough, balanced body; average clarity, almost no disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as anger, concentration, malevolence malice; style: cartoonish, authoritative; average recording, quiet background; genuineness 0.8/6; vocal-burst blend 1.6/10; 18.9s, PT.
portugal_15_1_153_2203439_2222336 · in -19.6 dBFS · gain -0.4 dB · eurospeech-02615
(contempt, concentration, disgust· measured, energised, slightly relaxed, cartoonish)Esta proposta de lei não só revela desrespeito pelos vários profissionais, ao tratar de forma igual setores com especificidades diferentes, como constitui uma inaceitável tentativa de controlo dos representantes das classes profissionais.
full caption & clip details
A middle-aged masculine voice; delivery is energised, measured, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, rough, thin; clear, almost no disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, fairly guarded; reads as contempt, concentration, disgust; style: cartoonish, storytelling; average recording, quiet background; genuineness 0.9/6; vocal-burst blend 1.7/10; 17.2s, PT.
portugal_15_1_153_2222336_2239584 · in -20.6 dBFS · gain +0.6 dB · eurospeech-02615
(disappointment, pride, shame·normal-paced, energised, neutral tension, cartoonish)Isto para não falar na audição das ordens, que foi um mero formalismo para cumprir calendário, uma vez que nada do que disseram foi tido em conta no documento que apresentaram.
full caption & clip details
A middle-aged masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, rough, thin; clear, almost no disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, fairly guarded; reads as disappointment, pride, shame; style: cartoonish, storytelling; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 2.5/10; 10.9s, PT.
portugal_15_1_153_2254192_2265104 · in -20.6 dBFS · gain +0.6 dB · eurospeech-02615
(bitterness, contempt, pride · normal-paced, energised, slightly relaxed, cartoonish)Sem dúvida que é importante eliminar obstáculos no desenvolvimento das atividades de serviços para promover o progresso económico e social, mas não é isso que esta proposta defende quando propõe a criação de um conselho de supervisão com personalidades de reconhecido mérito,
full caption & clip details
A middle-aged masculine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, almost no disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, fairly guarded; reads as bitterness, contempt, pride; style: cartoonish, authoritative; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 2.8/10; 16.7s, PT.
portugal_15_1_153_2265104_2281808 · in -21.8 dBFS · gain +1.8 dB · eurospeech-02615
The chain starts with Impatience and Irritability strongly present — 0.77, higher than 77 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.22.
It takes 2 clips to get there. Clip to clip the moves are +0.22 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.63 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.63 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 20 s · zh · emolia
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.573 before conversion and 0.626 after — it rose by 0.052. Neighbour-to-neighbour the worst pair went 0.573 → 0.626. (The earlier render, with segment 1 left raw, scores 0.578 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.222 in the original and +0.147 after conversion — 66 % of the delta retained. On the other named axis, Relief, -0.497 became -0.563.
Quality. Mean predicted overall quality across the segments went 2.77 → 2.94 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.573 → 0.626+0.052identity cos neighbours 0.573 → 0.626d_b rescored +0.222 → +0.147d_a rescored -0.497 → -0.563d_a mined -0.497d_b mined 0.222min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00076_S05918total 19.7schain gain +1.8 dBseam step 2.1 dBcrossfades 100 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an elderly feminine voice · neutral-toned, average recording, moderately variable, very wide pitch range
An elderly feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, slightly thin; somewhat unclear, some disfluency, very wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as relief, astonishment surprise, pleasure ecstasy; style: conversational, casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 3.9/10; 16.2s, ZH.
ZH_B00076_S05918_W000014 · in -24.2 dBFS · gain +4.2 dB · emolia-04040
(impatience and irritability, anger, bitterness·brisk, highly aroused, tense, cartoonish)不用了不用了,我从今以后,只换你。
full caption & clip details
A child masculine voice; delivery is highly aroused, brisk, tense, moderately variable; timbre is neutral-toned, slightly bright, very rough, thin; very clear, no disfluency, very wide pitch range, heavy breath; affect is elated, very dominant, guarded; reads as impatience and irritability, anger, bitterness; style: cartoonish, dramatic; average recording, some background noise; genuineness 3.1/6; vocal-burst blend 3.8/10; 3.7s, ZH.
ZH_B00076_S05918_W000015 · in -18.2 dBFS · gain -1.8 dB · emolia-04040
The chain starts with Doubt clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.28.
It takes 5 clips to get there. Clip to clip the moves are -0.03, then +0.14, then +0.07, then +0.09 — not a clean run: step 1 moves back the other way by 0.03 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.60 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.63 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.60, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 122 s · nl · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.568 before conversion and 0.884 after — it rose by 0.317. Neighbour-to-neighbour the worst pair went 0.558 → 0.877. (The earlier render, with segment 1 left raw, scores 0.718 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.274 in the original and +0.165 after conversion — 60 % of the delta retained. On the other named axis, Interest, -0.086 became -0.001.
Quality. Mean predicted overall quality across the segments went 2.69 → 3.27 (+0.58) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.568 → 0.884+0.317identity cos neighbours 0.558 → 0.877d_b rescored +0.274 → +0.165d_a rescored -0.086 → -0.001d_a mined -0.087d_b mined 0.275min_cos_consec (site) 0.6311min_cos_anchor (site) 0.6034dataset podcastlang nlspeaker 262338total 120.9schain gain +3.4 dBseam step 4.2 dBcrossfades 150/100/100/100 ms
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · quiet background, neutral tension
(interest, concentration, contemplation · brisk, energised, moderately variable, monologue)(ahem) Vos lagen gestoord. En natuurlijk niet in lijn met de boodschap van Christus. Alleen die mensen, ja, dat waren wel christenen. Ik vind het nou leuk. Die organisatie was zelfs christelijk. (ahem) (ahem) Maar dat rekenen we ook het christendom niet aan. Dus ik probeer altijd een beetje de voorbeelden erbij te halen vanuit andere tradities. En dat zie je ook bijvoorbeeld in het hindoeïsme nu. Ik wil je Hindva in India. waar echt echt de Hindoe-nationalisten te keer gaan als wilden tegen moslims, omdat ze moslims zijn.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, dark, fairly smooth, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as interest, concentration, contemplation; style: monologue, casual; below-average recording, quiet background; genuineness 3.9/6; vocal-burst blend 10.0/10; 26.6s, NL.
262338_00389808 · in -22.8 dBFS · gain +2.8 dB · podcast-05563
(concentration, contempt, interest ·normal-paced, normally alert, fairly steady, casual)sommige moslims te zien als (surprised gasp) de representatie van de islam. terwijl je dat bij andere godzien niet zo ziet. Ik denk dat je altijd gewoon naar de basis moet kijken. Het voorbeeld. In dit geval is dat natuurlijk gelukkig de profeetfeesten met hem. Ja, daar zie je (low mumble) prachtig gedrag, daar zie je geen gekke dingen. Je kan de contexten altijd uitleggen waar het (low mumble) onduidelijk is voor mensen. (ahem) En dat is de leidraad, dat is waar we naar moeten kijken.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, contempt, interest; style: casual, monologue; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 9.2/10; 23.6s, NL.
262338_00393532 · in -22.4 dBFS · gain +2.4 dB · podcast-05531
(sourness, relief, contemplation· normal-paced, normally alert, fairly steady, monologue)Dus als mensen die vraag stellen, ik probeer toch altijd een vergelijking te maken met andere tradities om het ze uit te leggen. En vervolgens te wijzen van van kijk, dit is wat de leerverkondigt. Dit is hoe we ons dienen te gedragen. En dat is waar je de godziens op mag (ahem) (low mumble) afrekenen.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, relief, contemplation; style: monologue, casual; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 7.9/10; 26.4s, NL.
262338_00395888 · in -22.9 dBFS · gain +2.9 dB · podcast-05552
(jealousy and envy, concentration·fast, very low-energy, fairly steady, casual)Ja, (low mumble) (low mumble) (ahem) ik (ahem) denk dat je gewoon uiteindelijk toch (low mumble) die beslissingen zou moeten nemen om te bekeren als dat echt jouw overtuiging is. (ahem) Ik denk ook dat die klopt dan. Maar (ahem)
full caption & clip details
A young adult masculine voice; delivery is very low-energy, fast, neutral tension, fairly steady; timbre is slightly cool, dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, concentration; style: casual, monologue; below-average recording, quiet background; genuineness 4.7/6; vocal-burst blend 10.0/10; 23.7s, NL.
262338_00398524 · in -23.3 dBFS · gain +3.3 dB · podcast-04143
(doubt, contemplation·normal-paced, very low-energy, fairly steady, casual)(low mumble) het is vaak ook je levenswandel, die op een gegeven moment mensen laat zien dat het wat minder spannend is dan zij denken. En dat het in sommige gevallen misschien zelfs inspirerend werkt. Ik merkte het ook in mijn omgeving, (low mumble) wat ik net al aangaf, mijn vrouw die (low mumble) tijdens de Roma gaf ze dan voor het eerst aan, ik vind je zo geduldig.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, thin; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, contemplation; style: casual, storytelling; below-average recording, quiet background; genuineness 5.2/6; vocal-burst blend 8.1/10; 21.2s, NL.
262338_00400888 · in -23.5 dBFS · gain +3.5 dB · podcast-04156
The chain starts with Doubt clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.35.
It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.15 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.83 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.83 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 39 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.612 before conversion and 0.480 after — it fell by 0.132. Neighbour-to-neighbour the worst pair went 0.582 → 0.590. (The earlier render, with segment 1 left raw, scores 0.383 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.352 in the original and +0.292 after conversion — 83 % of the delta retained, which is most of it. On the other named axis, Contemplation, -0.208 became -0.458.
Quality. Mean predicted overall quality across the segments went 2.47 → 3.02 (+0.55) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.612 → 0.480-0.132identity cos neighbours 0.582 → 0.590d_b rescored +0.352 → +0.292d_a rescored -0.208 → -0.458d_a mined -0.208d_b mined 0.352min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_JZ0qvzuEFyMtotal 38.6schain gain +2.1 dBseam step 1.7 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged feminine voice · slightly thin, average recording, quiet background, measured, frequent disfluency
(contemplation, hope enthusiasm optimism, pleasure ecstasy · very low-energy, relaxed, moderately variable, casual)It says, well we have had so many lifetimes and in all those lifetimes we have done so many things, good things, bad things, neutral things, eventually the results of those will come up like planting black seeds and white seeds,
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is slightly warm, slightly bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as contemplation, hope enthusiasm optimism, pleasure ecstasy; style: casual, ASMR; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.9/10; 17.6s, EN.
EN_JZ0qvzuEFyM_W000069 · in -17.2 dBFS · gain -2.8 dB · emolia-02229
(concentration, fear· very low-energy, neutral tension, moderately variable, casual)When the causes and conditions come, they will ripen. So we can't do much about that. Those seeds have already been sown. But what we can do is how we respond to those seeds as they come up, if they are weeds,
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as concentration, fear; style: casual, didactic; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 2.9/10; 18.0s, EN.
EN_JZ0qvzuEFyM_W000070 · in -21.3 dBFS · gain +1.3 dB · emolia-02229
(doubt·normally alert, slightly relaxed, fairly steady, casual)Or if they are good plants, we can make use of both.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as doubt; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 2.6/10; 3.4s, EN.
EN_JZ0qvzuEFyM_W000071 · in -17.1 dBFS · gain -2.9 dB · emolia-02229
The chain starts with style: dramatic (S_DRAM) around average — 0.53, higher than 53 % of clips in this corpus — and ends with it high at 0.78, higher than 78 % of clips in this corpus. That is a total rise of 0.25.
It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 40 s · da · podcast
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.947 before conversion and 0.913 after — it fell by 0.035. Neighbour-to-neighbour the worst pair went 0.947 → 0.913. (The earlier render, with segment 1 left raw, scores 0.866 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, S_DRAM — style: dramatic moved +0.254 in the original and +0.121 after conversion — 47 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 3.16 → 3.42 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.947 → 0.913-0.035identity cos neighbours 0.947 → 0.913d_b rescored +0.254 → +0.121d_a rescored +0.254 → +0.121d_a mined 0.249d_b mined 0.249min_cos_consec (site) 0.9467min_cos_anchor (site) 0.9467dataset podcastlang daspeaker 914933total 39.5schain gain +3.5 dBseam step 1.4 dBcrossfades 100 ms
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, moderately variable, average clarity, light breath
(intoxication altered states of consciousness · normal-paced, normally alert, neutral tension, casual)Og (ahem) Rising for Dorn har jeg prøvet og spillet, og så videre. (ahem) Men de andre to (ahem) af. (low mumble) Men prøv. Jeg synes, det er super super fedt. (ahem) At de gør det. Og det er jo også bare konkurrence til Microsoft. De kalder det play at home. (ahem) Hvad hedder det. Det er
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is positive, slightly submissive, slightly guarded; reads as intoxication altered states of consciousness; style: casual, conversational; average recording, some background noise; genuineness 6.0/6; vocal-burst blend 8.5/10; 16.6s, DA.
914933_00523567 · in -27.0 dBFS · gain +7.0 dB · podcast-03935
(elation, interest, hope enthusiasm optimism·brisk, energised, slightly relaxed, casual)Det er en ny ting. Det der er interessant ved det her. (ahem) Hvis jeg lige skal tage den kan skæld på. Det er det sådan lidt Gary fra IGN, der har pointeret det her. (low mumble) Fakt, jeg havde ikke selv lagt mærke til det. Men der har været rigtig stille på Twitter som sagt, i forhold til, altså det har været des ikke det andskyld. Det har været (low mumble) gamepass og Xbox, der har trendet helt vildt meget her de sidste par måneder. Og. (ahem)
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as elation, interest, hope enthusiasm optimism; style: casual, playful; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 9.6/10; 23.1s, DA.
914933_00525296 · in -29.9 dBFS · gain +9.9 dB · podcast-06401
The chain starts with Emotional Numbness below average — 0.32, lower than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.62.
It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.04, then +0.17, then +0.20 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.87 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.87 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 37 s · fr · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.813 before conversion and 0.744 after — it fell by 0.068. Neighbour-to-neighbour the worst pair went 0.740 → 0.663. (The earlier render, with segment 1 left raw, scores 0.712 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.620 in the original and +0.048 after conversion — 8 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Pleasure Ecstasy, -0.767 became +0.572.
Quality. Mean predicted overall quality across the segments went 2.89 → 3.04 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.813 → 0.744-0.068identity cos neighbours 0.740 → 0.663d_b rescored +0.620 → +0.048d_a rescored -0.767 → +0.572d_a mined -0.767d_b mined 0.620min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang frspeaker FR_KJr--UwglKItotal 35.7schain gain +0.6 dBseam step 0.8 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(fast, almost no disfluency, clear, authoritative)Cela dit, la news (ahem) primordiale, c'est bien entendu le championnat du monde de yoga sur poteaux, que nous avons tous attendu avec impatience.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.4/10; 6.5s, FR.
FR_KJr--UwglKI_W000000 · in -16.5 dBFS · gain -3.5 dB · emolia-02863
(measured, some disfluency, average clarity, monologue)Oh. Oh, qu'est-ce qu'il y a, oui.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 4.3/10; 13.8s, FR.
FR_KJr--UwglKI_W000001 · in -19.6 dBFS · gain -0.4 dB · emolia-02863
(awe, astonishment surprise·normal-paced, some disfluency, average clarity, conversational)Voilà. Des acrobaties incroyables qui laissent vraiment un rêve-œuvre.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as awe, astonishment surprise; style: conversational, playful; average recording, no background noise; genuineness 4.2/6; vocal-burst blend 0.0/10; 3.6s, FR.
FR_KJr--UwglKI_W000002 · in -16.7 dBFS · gain -3.3 dB · emolia-02863
(fast, little disfluency, clear, authoritative)Alors, on rappelle que ce sport qui attire de plus en plus de pratiquants reste (ahem) néanmoins très risqué.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.5/10; 5.1s, FR.
FR_KJr--UwglKI_W000003 · in -17.9 dBFS · gain -2.1 dB · emolia-02863
(emotional numbness· fast, little disfluency, clear, authoritative)Eh oui, car une perte de slip ou une blessure par écharde peut se produire à tout moment. Donc si vous voulez le pratiquer depuis chez vous, restez prudents.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: authoritative, monologue; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 2.8/10; 7.5s, FR.
FR_KJr--UwglKI_W000004 · in -17.8 dBFS · gain -2.2 dB · emolia-02863
The chain starts with Longing strongly present — 0.76, higher than 76 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.22.
It takes 5 clips to get there. Clip to clip the moves are +0.18, then +0.05, then -0.05, then +0.05 — not a clean run: step 3 moves back the other way by 0.05 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.18 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.28 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.18, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 40 s · en · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.223 before conversion and 0.687 after — it rose by 0.464. Neighbour-to-neighbour the worst pair went 0.223 → 0.586. (The earlier render, with segment 1 left raw, scores 0.497 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.225 in the original and +0.283 after conversion — 126 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contemplation, -0.596 became -0.456.
Quality. Mean predicted overall quality across the segments went 2.53 → 2.95 (+0.42) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.223 → 0.687+0.464identity cos neighbours 0.223 → 0.586d_b rescored +0.225 → +0.283d_a rescored -0.596 → -0.456d_a mined -0.606d_b mined 0.224min_cos_consec (site) 0.2758min_cos_anchor (site) 0.1798dataset podcastlang enspeaker 888667total 39.0schain gain +3.5 dBseam step 1.6 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · normally alert, moderately variable
(contemplation, doubt, infatuation · normal-paced, relaxed, frequent disfluency, casual)camper. (low mumble) Um I yeah, I think like we also try to put make it so that you can move. You can't
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as contemplation, doubt, infatuation; style: casual, conversational; good recording, quiet background; genuineness 3.7/6; vocal-burst blend 3.5/10; 7.5s, EN.
888667_00172592 · in -31.1 dBFS · gain +11.1 dB · podcast-03824
(pleasure ecstasy, contentment, relief·brisk, neutral tension, some disfluency, casual)daughter's elementary school and do reading. So I'd go into her classroom. So one year I went as Captain Books. I had my pirate outfit on and I had my books. That was really fun. And then one year my youngest and I went to Comic Con.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, slightly rough, slightly thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as pleasure ecstasy, contentment, relief; style: casual, playful; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 8.8/10; 15.6s, EN.
888667_00174400 · in -35.9 dBFS · gain +15.9 dB · podcast-02661
(amusement, embarrassment, longing·normal-paced, relaxed, some disfluency, casual)And I got a little (low mumble) um headband. I still have it on my desk. It was like (ahem) uh it was a homemade headband, but it had like horns and then it had like
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; somewhat unclear, some disfluency, moderate pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as amusement, embarrassment, longing; style: casual, conversational; below-average recording, quiet background; genuineness 5.3/6; vocal-burst blend 7.5/10; 7.5s, EN.
888667_00176040 · in -35.7 dBFS · gain +15.7 dB · podcast-03793
(pleasure ecstasy, embarrassment, infatuation· normal-paced, relaxed, some disfluency, casual)One other question I did that was really fun that I really liked. I don't know.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is slightly cool, dark, fairly smooth, thin; slurred, some disfluency, wide pitch range, heavy breath; affect is mildly negative, slightly submissive, neutral openness; reads as pleasure ecstasy, embarrassment, infatuation; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 5.5/6; vocal-burst blend 3.4/10; 4.4s, EN.
888667_00178312 · in -36.7 dBFS · gain +16.7 dB · podcast-00145
(longing, embarrassment, amusement· normal-paced, neutral tension, some disfluency, casual)was Max One Year from Where the Wild Things Are. I think that was my favorite now that we're talking.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as longing, embarrassment, amusement; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.0/6; vocal-burst blend 3.1/10; 4.7s, EN.
888667_00178920 · in -29.7 dBFS · gain +9.7 dB · podcast-02643
The chain starts with Doubt around average — 0.55, higher than 55 % of clips in this corpus — and ends with it strongly present at 0.78, higher than 78 % of clips in this corpus. That is a total rise of 0.23.
It takes 3 clips to get there. Clip to clip the moves are +0.18, then +0.05 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of -0.10 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of -0.10 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 31 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.042 before conversion and 0.310 after — it rose by 0.352. Neighbour-to-neighbour the worst pair went -0.042 → 0.524. (The earlier render, with segment 1 left raw, scores 0.291 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.231 in the original and +0.021 after conversion — 9 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Fear, -0.519 became -0.557.
Quality. Mean predicted overall quality across the segments went 2.69 → 2.96 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 -0.042 → 0.310+0.352identity cos neighbours -0.042 → 0.524d_b rescored +0.231 → +0.021d_a rescored -0.519 → -0.557d_a mined -0.519d_b mined 0.230min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_p64KLCRGpVItotal 30.1schain gain +1.5 dBseam step 1.8 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, average recording, quiet background, slightly relaxed, some disfluency, moderate pitch range
(fear · normal-paced, subdued, fairly steady, casual)I don't make those changes, but the reality is, (low mumble) uhm, we're gonna get a new system and you either get what you want or you're gonna be forced into something. And sometimes it's not easy. Yeah, they definitely know what they want because they know their job better than anybody. The reality is extracting that from them is sometimes a long and painful process until you put it on them to say, okay, here's what you're gonna get. And then they'll be very (low mumble) willing to give you more feedback.
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as fear; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 8.8/10; 21.7s, EN.
EN_p64KLCRGpVI_W000083 · in -21.7 dBFS · gain +1.7 dB · emolia-01629
(normal-paced, normally alert, moderately variable, conversational)Yeah, I, I mean I think we just need better tools. We don't have good requirements engineering tools.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: conversational, casual; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 2.8/10; 5.3s, EN.
EN_p64KLCRGpVI_W000084 · in -21.7 dBFS · gain +1.7 dB · emolia-01629
(measured, normally alert, fairly steady, casual)But better tool, that doesn't mean that we shouldn't do it.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: casual, formal; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 2.6/10; 3.5s, EN.
EN_p64KLCRGpVI_W000085 · in -19.1 dBFS · gain -0.9 dB · emolia-01629
The chain starts with Impatience and Irritability strongly present — 0.76, higher than 76 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.22.
It takes 3 clips to get there. Clip to clip the moves are +0.08, then +0.14 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.78 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 46 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.677 before conversion and 0.659 after — it fell by 0.018. Neighbour-to-neighbour the worst pair went 0.708 → 0.682. (The earlier render, with segment 1 left raw, scores 0.546 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.211 in the original and +0.091 after conversion — 43 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Amusement, -0.149 became -0.147.
Quality. Mean predicted overall quality across the segments went 2.89 → 3.15 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.677 → 0.659-0.018identity cos neighbours 0.708 → 0.682d_b rescored +0.211 → +0.091d_a rescored -0.149 → -0.147d_a mined -0.148d_b mined 0.216min_cos_consec (site) 0.7794min_cos_anchor (site) 0.8366dataset podcastlang enspeaker 498239total 44.9schain gain +3.3 dBseam step 1.5 dBcrossfades 100/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, wide pitch range, light breath
(amusement, pleasure ecstasy, teasing · brisk, normally alert, neutral tension, casual)like that extra 10. And like I've done an hour on the road before. I'm an hour for me is a stretch at this point. Like I can do a comfortable 40, comfortable 45, but still like just taking that 10 minute jump at Yucky's. I was a little nervous. I'm not gonna lie, I was a little nervous. Uh, (low mumble) and that's why I only charge you $500 to come to Ottawa and a hotel room and not $1,500. Cause uh anyway, so I (low mumble) uh
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as amusement, pleasure ecstasy, teasing; style: casual, conversational; average recording, quiet background; genuineness 5.5/6; vocal-burst blend 9.7/10; 23.3s, EN.
498239_00086648 · in -24.2 dBFS · gain +4.2 dB · podcast-06322
(triumph, sexual lust, pleasure ecstasy ·normal-paced, energised, neutral tension, casual)(low mumble) was was gonna go (low mumble) um next, and there was a girl before me that smashed really hard, and when that makes you a little even more nervous, like, oh man, I gotta follow somebody that did great, and you know what? I fucking killed it. I smashed like one of the best sets I've had, like from beginning to end, really good. Amazing opening, great closing.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as triumph, sexual lust, pleasure ecstasy; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.9/6; vocal-burst blend 10.0/10; 17.6s, EN.
498239_00088976 · in -23.8 dBFS · gain +3.8 dB · podcast-06305
(impatience and irritability, contempt, anger·measured, normally alert, slightly relaxed, casual)People were gassing me up after the show. (low mumble) Uh, these
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as impatience and irritability, contempt, anger; style: casual, conversational; average recording, no background noise; genuineness 3.1/6; vocal-burst blend 0.0/10; 4.4s, EN.
498239_00090728 · in -23.2 dBFS · gain +3.2 dB · podcast-04519
The chain starts with Concentration clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.21.
It takes 2 clips to get there. Clip to clip the moves are +0.21 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.91 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.91 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 32 s · en · emolia
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.869 before conversion and 0.818 after — it fell by 0.051. Neighbour-to-neighbour the worst pair went 0.869 → 0.818. (The earlier render, with segment 1 left raw, scores 0.564 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.211 in the original and +0.202 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Fear, -0.505 became -0.143.
Quality. Mean predicted overall quality across the segments went 3.10 → 3.26 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.869 → 0.818-0.051identity cos neighbours 0.869 → 0.818d_b rescored +0.211 → +0.202d_a rescored -0.505 → -0.143d_a mined -0.505d_b mined 0.210min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_UTnqPEGiwMEtotal 31.2schain gain +2.3 dBseam step 0.1 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fear · fairly steady, formal, newsreading)It may be accompanied with elaborate prayers, other rites such as charity or visit to a temple, sometimes observed during festivals or with sanskara
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 9.9s, EN.
EN_UTnqPEGiwME_W000015 · in -15.0 dBFS · gain -5.0 dB · emolia-01602
(concentration·steady, newsreading, formal)The Puranas link the practice to the empowering concept of Shakti of a woman, while the Dharmazastras link the practice to one possible form of penance through the concept of prayaschita for both men and women.Avrata is a personal practice, typically involves no priest, but may involve personal prayer, chanting, reading of spiritual texts, social get-together of friends and family, or silent meditation.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 21.6s, EN.
EN_UTnqPEGiwME_W000017 · in -15.4 dBFS · gain -4.5 dB · emolia-01602
The chain starts with audible breath / respiration (RESP) around average — 0.55, higher than 55 % of clips in this corpus — and works its way down to low at 0.13, lower than 87 % of clips in this corpus. That is a total fall of 0.42.
It takes 4 clips to get there. Clip to clip the moves are -0.10, then -0.13, then -0.20 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.90 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 86 s · de · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.848 before conversion and 0.907 after — it rose by 0.060. Neighbour-to-neighbour the worst pair went 0.910 → 0.929. (The earlier render, with segment 1 left raw, scores 0.842 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, RESP — audible breath / respiration moved -0.437 in the original and -0.424 after conversion — 97 % of the delta retained, which is essentially all of it.
Quality. Mean predicted overall quality across the segments went 3.14 → 3.33 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.848 → 0.907+0.060identity cos neighbours 0.910 → 0.929d_b rescored -0.437 → -0.424d_a rescored -0.437 → -0.424d_a mined -0.425d_b mined -0.425min_cos_consec (site) 0.9004min_cos_anchor (site) 0.8862dataset podcastlang despeaker 524119total 84.8schain gain +5.0 dBseam step 1.7 dBcrossfades 150/150/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, average clarity, moderate pitch range, light breath
(relief, thankfulness gratitude, fear · normal-paced, normally alert, neutral tension, conversational)Und das ist ganz klar bei mir (low mumble) auch ein ganz wichtiger Punkt der Auftragsklärung, dass ich dann sage: Hey, was ist denn euer Wunsch? Also, ich habe ja einmal den externen Wunsch, aber ihr als Team, ihr seid ja,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as relief, thankfulness gratitude, fear; style: conversational, casual; good recording, quiet background; genuineness 4.5/6; vocal-burst blend 1.8/10; 13.4s, DE.
524119_00206967 · in -29.3 dBFS · gain +9.3 dB · podcast-01065
(contemplation, concentration, bitterness·slow, very low-energy, relaxed, monologue)und das ist, glaube ich, das ganz Besondere, ihr seid Individuen mit individuellen Ressourcen, mit individuellen Potenzialen, mit individuellen Fragen und Problemen. Und das ist natürlich die Person an sich, und jetzt aber betrachte ich diesen systemischen Ansatz, (ahem) dass ich sage,
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contemplation, concentration, bitterness; style: monologue, casual; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 3.2/10; 20.9s, DE.
524119_00208312 · in -30.4 dBFS · gain +10.4 dB · podcast-01069
(sourness, infatuation, sexual lust·normal-paced, normally alert, slightly relaxed, storytelling)ihr seid aber in einer Konstellation unterwegs, in einer Abhängigkeit, in einer Motivation auch zueinander, die einen direkten Einfluss auf euch habt. Also jeder kennt es, ne? Es gibt die eine Person in der Gruppe, die kommt immer freudestrahlend rein. Und das macht was mit diesem Team. Wenn diese Person krank ist oder nicht da ist, auf Urlaub ist, dann geht es dem Team anders.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as sourness, infatuation, sexual lust; style: storytelling, monologue; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 2.8/10; 23.0s, DE.
524119_00210400 · in -27.6 dBFS · gain +7.6 dB · podcast-06124
(sourness, triumph, jealousy and envy· normal-paced, normally alert, slightly relaxed, conversational)Und es gibt diese eine Person, die immer redet und immer über irgendwelche Sachen redet, die gerade keinen eigentlich interessieren und immer vom Thema abschweifen. Und ich habe schon (breathy giggle) Teams erlebt, wo dann wirklich gesagt wurde: Oh, jetzt können wir endlich mal produktiv arbeiten, die Person ist (low mumble) eine Woche im Urlaub. Und das ist natürlich auch in gewisser Weise toxisch für das Team. Und wichtig ist aber zu sehen, wann sind welche Ressourcen sinnvoll. Und jetzt als allerwichtigste Offenheit.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as sourness, triumph, jealousy and envy; style: conversational, casual; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 7.3/10; 28.0s, DE.
524119_00212696 · in -27.1 dBFS · gain +7.2 dB · podcast-06120