Manifest tier. emotion, rule B1, T=0.6, step cap 0.25. Population 100,844 chains (1,224 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 82,936.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words
are a real non-speech sound, printed where it happens.
The full generated caption for any clip is under “full caption & clip
details”. Its perceived-gender and background-noise clauses were re-rendered from
the numeric buckets, because the versions stored in the corpus index had those two
ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
The tags and captions describe the ORIGINAL clips — they are the
corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the
converted audio are the ones printed in each card's own paragraph and chips.
This chain comes from the one-sided rule: only Astonishment Surprise had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Astonishment Surprise below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.64.
Nothing was asked of the other axis, and in fact Teasing drifts down from 0.99 to 0.73 (-0.26), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.16, then +0.24, then +0.23 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 34 s · snippets
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.145 before conversion and 0.543 after — it rose by 0.398. Neighbour-to-neighbour the worst pair went 0.230 → 0.543. (The earlier render, with segment 1 left raw, scores 0.306 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.637 in the original and +0.667 after conversion — 105 % of the delta retained, which is essentially all of it. On the other named axis, Teasing, -0.264 became -0.082.
Quality. Mean predicted overall quality across the segments went 2.80 → 3.03 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.145 → 0.543+0.398identity cos neighbours 0.230 → 0.543d_b rescored +0.637 → +0.667d_a rescored -0.264 → -0.082d_a mined -0.261d_b mined 0.637min_cos_consec (site) —min_cos_anchor (site) —dataset snippetslang ?speaker batch79_part1_batch79_parttotal 33.1schain gain +3.3 dBseam step 1.3 dBcrossfades 150/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-bright, average clarity
(teasing, sexual lust, impatience and irritability · brisk, energised, slightly tense, casual)They're going restrained with little to eat and just tends to sleep. You're restrained? I am.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is slightly cool, neutral-bright, very rough, thin; average clarity, almost no disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, fairly guarded; reads as teasing, sexual lust, impatience and irritability; style: casual, storytelling; average recording, some background noise; genuineness 3.2/6; vocal-burst blend 1.4/10; 4.4s.
batch79_part1_batch79_part1_chunk_1703_1_1521767 · in -22.7 dBFS · gain +2.7 dB · snippets-01294
(normal-paced, normally alert, slightly relaxed, casual)economy plane ticket would be the least of their worries, as the guest would soon land.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; no dominant emotion; style: casual, storytelling; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 3.7/10; 4.3s.
batch79_part1_batch79_part1_chunk_1703_1_1521800 · in -24.4 dBFS · gain +4.4 dB · snippets-01294
(interest, hope enthusiasm optimism, elation· normal-paced, normally alert, slightly relaxed, casual)We're going to throw a festival yeah. Fire team said they wanted to do a music festival in the Bahamas, but they wanted to do it in only six months. Billy's plan was to promote the fire app with his very own music festival. And although Billy had no experience organizing such an event, what he did have was eight figures of venture capital.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, hope enthusiasm optimism, elation; style: casual, conversational; good recording, no background noise; genuineness 2.9/6; vocal-burst blend 6.2/10; 16.9s.
batch79_part1_batch79_part1_chunk_1703_1_1521909 · in -24.5 dBFS · gain +4.5 dB · snippets-01294
(astonishment surprise, fear, distress· normal-paced, normally alert, neutral tension, casual)But this still wasn't even close to enough. Billy had said, listen, we're into this for 25 million bucks. I'm like, oh my gosh, Billy, sheesh.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as astonishment surprise, fear, distress; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 3.9/6; vocal-burst blend 3.0/10; 7.9s.
batch79_part1_batch79_part1_chunk_1703_1_1522202 · in -26.6 dBFS · gain +6.6 dB · snippets-01294
Jealousy and Envy ↑ (unconstrained axis: Sexual Lust)identity +0.01emotion 121 % emotion__B1__T0.60__C0.25__INTERNAL · #2
This chain comes from the one-sided rule: only Jealousy and Envy had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Jealousy and Envy below average — 0.26, lower than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.65.
Nothing was asked of the other axis, and in fact Sexual Lust drifts down from 0.98 to 0.33 (-0.65), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.10, then +0.13, then +0.17 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 45 s · ko · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.650 before conversion and 0.657 after — it rose by 0.007. Neighbour-to-neighbour the worst pair went 0.847 → 0.771. (The earlier render, with segment 1 left raw, scores 0.636 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Jealousy and Envy moved +0.651 in the original and +0.786 after conversion — 121 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Sexual Lust, -0.647 became -0.737.
Quality. Mean predicted overall quality across the segments went 2.91 → 3.06 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.650 → 0.657+0.007identity cos neighbours 0.847 → 0.771d_b rescored +0.651 → +0.786d_a rescored -0.647 → -0.737d_a mined -0.647d_b mined 0.651min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang kospeaker KO_uT3tqbL3Ntytotal 43.5schain gain -0.4 dBseam step 1.7 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, normally alert, moderate pitch range, light breath
(sexual lust, doubt, sourness · normal-paced, slightly relaxed, fairly steady, storytelling)개면년은 2023년도를 이야기를 하는 상황이고,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as sexual lust, doubt, sourness; style: storytelling, casual; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 4.2/10; 4.0s, KO.
KO_uT3tqbL3Nty_W000090 · in -13.2 dBFS · gain -6.8 dB · emolia-03170
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: formal, casual; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 3.3/10; 4.3s, KO.
KO_uT3tqbL3Nty_W000091 · in -15.8 dBFS · gain -4.2 dB · emolia-03170
(normal-paced, slightly relaxed, fairly steady, monologue)이 임계수에 당하는 관인을 쓰기가 사실 쉽지가 않아요. 지는 땅이거든요. 반죽이 잘 된 땅인데, 땅 속에 묻혀있는 지, 이,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 1.8/6; vocal-burst blend 1.4/10; 9.2s, KO.
KO_uT3tqbL3Nty_W000092 · in -16.8 dBFS · gain -3.2 dB · emolia-03170
(pride, infatuation, longing·measured, relaxed, moderately variable, whispered)이렇게 개면연으로 두르게 된다면, 진토가 목곡을 이루게 되면서, 진토가 결국 희생을 하게 돼요. 즉, 식신이 희생을 당하게 되는 말 그대로, 나의 재능이나, 또, 본인의 어떤 아이디어라던가, 자신의 재능이 그대로,
full caption & clip details
An elderly feminine voice; delivery is normally alert, measured, relaxed, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, slightly thin; clear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as pride, infatuation, longing; style: whispered, casual; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 5.0/10; 19.8s, KO.
KO_uT3tqbL3Nty_W000093 · in -18.0 dBFS · gain -2.0 dB · emolia-03170
(jealousy and envy·normal-paced, slightly relaxed, fairly steady, casual)당장에 합격을 기대를 하신다면 좀 많이 버릴 수가 있습니다. 시간이 많이 걸리고 기다려야 되는.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as jealousy and envy; style: casual, conversational; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 3.8/10; 7.0s, KO.
KO_uT3tqbL3Nty_W000094 · in -19.2 dBFS · gain -0.8 dB · emolia-03170
This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Emotional Numbness barely there — 0.13, lower than 87 % of clips in this corpus — and ends with it clearly present at 0.73, higher than 74 % of clips in this corpus. That is a total rise of 0.61.
Nothing was asked of the other axis, and in fact Infatuation drifts down from 0.89 to 0.30 (-0.59), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.07, then +0.13, then +0.21 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.90 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.90 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 26 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.821 before conversion and 0.743 after — it fell by 0.078. Neighbour-to-neighbour the worst pair went 0.814 → 0.743. (The earlier render, with segment 1 left raw, scores 0.685 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.610 in the original and +0.471 after conversion — 77 % of the delta retained, which is most of it. On the other named axis, Infatuation, -0.587 became -0.392.
Quality. Mean predicted overall quality across the segments went 2.92 → 3.08 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.821 → 0.743-0.078identity cos neighbours 0.814 → 0.743d_b rescored +0.610 → +0.471d_a rescored -0.587 → -0.392d_a mined -0.587d_b mined 0.610min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00024_S09660total 24.4schain gain +0.1 dBseam step 2.1 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, no disfluency, light breath
This chain comes from the one-sided rule: only Sexual Lust had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Sexual Lust below average — 0.33, lower than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.62.
Nothing was asked of the other axis, and in fact Disappointment drifts down from 0.97 to 0.39 (-0.58), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.15, then +0.10, then +0.14, then +0.23 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.75 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.75 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 59 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.707 before conversion and 0.695 after — it fell by 0.012. Neighbour-to-neighbour the worst pair went 0.749 → 0.678. (The earlier render, with segment 1 left raw, scores 0.577 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sexual Lust moved +0.625 in the original and +0.610 after conversion — 98 % of the delta retained, which is essentially all of it. On the other named axis, Disappointment, -0.579 became -0.588.
Quality. Mean predicted overall quality across the segments went 2.91 → 3.15 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.707 → 0.695-0.012identity cos neighbours 0.749 → 0.678d_b rescored +0.625 → +0.610d_a rescored -0.579 → -0.588d_a mined -0.579d_b mined 0.625min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B-0aIYt4Sfctotal 57.7schain gain +1.5 dBseam step 4.2 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · fairly smooth, average clarity
(disappointment, concentration, impatience and irritability · slow, very low-energy, slightly relaxed, casual)But the problem with the one on the left, which is data collected by people logging issues into their apps, is that you are subsetting who is going to take part in that from the start.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as disappointment, concentration, impatience and irritability; style: casual, monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.6/10; 16.9s, EN.
EN_B-0aIYt4Sfc_W000064 · in -19.9 dBFS · gain -0.1 dB · emolia-02627
(normal-paced, normally alert, slightly relaxed, monologue)Not only does a person, (low mumble) uhm, need a phone or an internet connection or an app, that person has to know about the initiative
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: monologue, casual; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 1.3/10; 8.3s, EN.
EN_B-0aIYt4Sfc_W000065 · in -19.4 dBFS · gain -0.6 dB · emolia-02627
(concentration·slow, very low-energy, slightly relaxed, casual)Take the time to actually care about the initiative and think that their government is going to, uh, (low mumble) effectively respond to their issue. (low mumble) Uhm, and,
full caption & clip details
A young adult somewhat feminine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as concentration; style: casual, monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 1.0/10; 12.0s, EN.
EN_B-0aIYt4Sfc_W000066 · in -20.6 dBFS · gain +0.6 dB · emolia-02627
(interest, pleasure ecstasy, elation·brisk, normally alert, neutral tension, casual)I live there and I can tell you that it is like a walker's paradise. I can also tell you that Capitol Hill has probably like
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as interest, pleasure ecstasy, elation; style: casual, conversational; good recording, no background noise; genuineness 4.1/6; vocal-burst blend 8.7/10; 15.5s, EN.
EN_B-0aIYt4Sfc_W000069 · in -20.4 dBFS · gain +0.4 dB · emolia-02627
(sexual lust, fear, shame·normal-paced, normally alert, slightly relaxed, casual)Like their PTAs are probably the most involved of like anywhere in the city.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as sexual lust, fear, shame; style: casual, monologue; good recording, no background noise; genuineness 2.4/6; vocal-burst blend 1.8/10; 5.8s, EN.
EN_B-0aIYt4Sfc_W000070 · in -15.9 dBFS · gain -4.1 dB · emolia-02627
This chain comes from the one-sided rule: only Embarrassment had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Embarrassment below average — 0.32, lower than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.66.
Nothing was asked of the other axis, and in fact Concentration drifts down from 1.00 to 0.81 (-0.19), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.18, then +0.25, then +0.24 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.18 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.13 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.18, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 76 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.208 before conversion and 0.691 after — it rose by 0.483. Neighbour-to-neighbour the worst pair went 0.155 → 0.604. (The earlier render, with segment 1 left raw, scores 0.407 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.651 in the original and +0.407 after conversion — 63 % of the delta retained. On the other named axis, Concentration, -0.186 became -0.417.
Quality. Mean predicted overall quality across the segments went 2.83 → 3.25 (+0.42) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.208 → 0.691+0.483identity cos neighbours 0.155 → 0.604d_b rescored +0.651 → +0.407d_a rescored -0.186 → -0.417d_a mined -0.185d_b mined 0.651min_cos_consec (site) 0.1262min_cos_anchor (site) 0.1821dataset podcastlang enspeaker 286836total 75.3schain gain +2.3 dBseam step 1.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-toned, balanced body, quiet background, fairly steady
(concentration, contemplation, contentment · normal-paced, subdued, slightly relaxed, monologue)resolve conflicts without resorting to bloodshed. And this is something he comes to again and again in the book. And it's something he very much puts out there as his own role, (low mumble) something that he can perform again as a scientist, someone who has this respected genealogy of helping to mediate dispute. So we see a lot of examples in this in the book. And there are other examples that come up in other sources that talk about this period as well that really highlight
full caption & clip details
A middle-aged masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as concentration, contemplation, contentment; style: monologue, casual; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 5.3/10; 28.0s, EN.
286836_00067311 · in -20.0 dBFS · gain +0.0 dB · podcast-06030
(measured, normally alert, slightly relaxed, monologue)again, really illustrating the sort of broader role of the Imam beyond his more strictly defined role as
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, authoritative; good recording, quiet background; genuineness 2.2/6; vocal-burst blend 0.8/10; 6.7s, EN.
286836_00070224 · in -16.6 dBFS · gain -3.4 dB · podcast-03709
(shame, contemplation, concentration· measured, very low-energy, slightly relaxed, whispered)Ismaili Imam. So I want to talk here about the Ismaili community, and specifically in 2009 that you became an Ismaili. Can you share a little bit about your story specifically how and why you chose to do this, and what particularly motivated or inspired you?
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as shame, contemplation, concentration; style: whispered, ASMR; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 5.3/10; 20.1s, EN.
286836_00070904 · in -21.8 dBFS · gain +1.8 dB · podcast-05774
(embarrassment, contemplation, contentment· measured, very low-energy, neutral tension, conversational)(ahem) but to to make a long story short, I I had been studying about (low mumble) Islam uh (low mumble) really since I was in my later years in high school. (low mumble) Uh I was raised in a Roman Catholic family, but you know, was always encouraged to to study to learn about other traditions. And I think one of the questions that drove me from early on was I would.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as embarrassment, contemplation, contentment; style: conversational, casual; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 4.3/10; 21.1s, EN.
286836_00074080 · in -18.4 dBFS · gain -1.6 dB · podcast-05780
This chain comes from the one-sided rule: only Malevolence Malice had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Malevolence Malice below average — 0.25, lower than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.66.
Nothing was asked of the other axis, and in fact Fear drifts down from 0.78 to 0.17 (-0.62), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.00, then +0.24, then +0.21 — a plateau around step 2, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.89 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.89 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 35 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.858 before conversion and 0.865 after — it rose by 0.007. Neighbour-to-neighbour the worst pair went 0.858 → 0.865. (The earlier render, with segment 1 left raw, scores 0.802 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Malevolence Malice moved +0.655 in the original and +0.710 after conversion — 108 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Fear, -0.618 became -0.499.
Quality. Mean predicted overall quality across the segments went 2.87 → 3.13 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.858 → 0.865+0.007identity cos neighbours 0.858 → 0.865d_b rescored +0.655 → +0.710d_a rescored -0.618 → -0.499d_a mined -0.618d_b mined 0.655min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00004_S03819total 33.3schain gain +2.9 dBseam step 0.5 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady
(normal-paced, almost no disfluency, clear, authoritative)这是英国脱离欧盟的第一大基本原因。第二大原因就是移民和劳工问题。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, formal; average recording, no background noise; genuineness 1.6/6; vocal-burst blend 0.7/10; 6.2s, ZH.
ZH_B00004_S03819_W000180 · in -20.3 dBFS · gain +0.3 dB · emolia-03321
(normal-paced, little disfluency, clear, authoritative)(ahem) 呃,英国本身呢在福利支出方面不是欧元区最慷慨。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 1.4/10; 4.6s, ZH.
ZH_B00004_S03819_W000181 · in -20.7 dBFS · gain +0.7 dB · emolia-03321
(brisk, almost no disfluency, clear, authoritative)也就是福利社保不是最优厚的,但是相比于东欧的一些国家还是比较慷慨的。
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, dramatic; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 3.2/10; 6.5s, ZH.
ZH_B00004_S03819_W000182 · in -20.5 dBFS · gain +0.5 dB · emolia-03321
(normal-paced, some disfluency, average clarity, authoritative)比如像失业相关的补贴呀,儿童养老、公共的医疗、住房、教育、残疾方面的补贴。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, monologue; average recording, no background noise; genuineness 2.0/6; vocal-burst blend 2.7/10; 6.9s, ZH.
ZH_B00004_S03819_W000183 · in -19.8 dBFS · gain -0.2 dB · emolia-03321
(malevolence malice· normal-paced, some disfluency, clear, authoritative)那么在英国,有一些家庭,什么工作都不做,一年可以拿到各种福利的补贴。两到三万英镑这个水平在英国就是个平均水平。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as malevolence malice; style: authoritative, didactic; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 3.4/10; 9.9s, ZH.
ZH_B00004_S03819_W000184 · in -18.2 dBFS · gain -1.8 dB · emolia-03321
This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Concentration barely there — 0.21, lower than 79 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.65.
Nothing was asked of the other axis, and in fact Astonishment Surprise drifts down from 0.97 to 0.36 (-0.61), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.05, then +0.16, then +0.23, then +0.21 — a plateau around step 1, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.85 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.85 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 36 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.792 before conversion and 0.768 after — it fell by 0.024. Neighbour-to-neighbour the worst pair went 0.701 → 0.755. (The earlier render, with segment 1 left raw, scores 0.730 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.651 in the original and +0.597 after conversion — 92 % of the delta retained, which is essentially all of it. On the other named axis, Astonishment Surprise, -0.610 became -0.769.
Quality. Mean predicted overall quality across the segments went 2.78 → 2.85 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.792 → 0.768-0.024identity cos neighbours 0.701 → 0.755d_b rescored +0.651 → +0.597d_a rescored -0.610 → -0.769d_a mined -0.610d_b mined 0.651min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00048_S06108total 35.1schain gain +2.4 dBseam step 1.3 dBcrossfades 100/100/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · fairly smooth, balanced body, good recording, no background noise, slightly relaxed, clear
(astonishment surprise, embarrassment, disappointment · normal-paced, normally alert, fairly steady, conversational)I think you've missed the boat on that application. They've already started interviewing candidates.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as astonishment surprise, embarrassment, disappointment; style: conversational, formal; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 1.9/10; 5.1s, EN.
EN_B00048_S06108_W000043 · in -26.8 dBFS · gain +6.8 dB · emolia-01182
(contentment, affection, pleasure ecstasy· normal-paced, energised, moderately variable, casual)Number 15 is a lovely one. It is a piece of cake, a piece of cake. This means really easy. That pop quiz was a piece of cake.
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; clear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as contentment, affection, pleasure ecstasy; style: casual, playful; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 0.1/10; 10.8s, EN.
EN_B00048_S06108_W000044 · in -24.9 dBFS · gain +4.9 dB · emolia-01182
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as sexual lust; style: playful, storytelling; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 0.0/10; 4.5s, EN.
EN_B00048_S06108_W000045 · in -25.3 dBFS · gain +5.3 dB · emolia-01182
(impatience and irritability·brisk, normally alert, moderately variable, dramatic)For example, I think you need to pull yourself together and stop stressing about the presentation.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, little disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as impatience and irritability; style: dramatic, casual; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 1.1/10; 5.0s, EN.
EN_B00048_S06108_W000046 · in -25.8 dBFS · gain +5.8 dB · emolia-01182
(normal-paced, normally alert, steady, dramatic)Number 17 is to sit or to be on the fence. To sit on the fence. To be on the fence. This means to stay neutral and to not take sides.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: dramatic, casual; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.3/10; 10.3s, EN.
EN_B00048_S06108_W000047 · in -26.0 dBFS · gain +6.0 dB · emolia-01182
This chain comes from the one-sided rule: only Shame had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Shame barely there — 0.25, lower than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.66.
Nothing was asked of the other axis, and in fact Contemplation drifts down from 0.87 to 0.79 (-0.08), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.22, then +0.25, then +0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.82 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.82 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 48 s · fr · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.492 before conversion and 0.437 after — it fell by 0.055. Neighbour-to-neighbour the worst pair went 0.424 → 0.436. (The earlier render, with segment 1 left raw, scores 0.353 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.660 in the original and +0.353 after conversion — 53 % of the delta retained. On the other named axis, Contemplation, -0.083 became -0.017.
Quality. Mean predicted overall quality across the segments went 2.90 → 3.10 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.492 → 0.437-0.055identity cos neighbours 0.424 → 0.436d_b rescored +0.660 → +0.353d_a rescored -0.083 → -0.017d_a mined -0.082d_b mined 0.660min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang frspeaker FR_OjAMb4folEItotal 47.0schain gain +3.6 dBseam step 2.6 dBcrossfades 150/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, slightly relaxed
(measured, normally alert, fairly steady, narration)la consommation énergétique pour réaliser
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, very full; very clear, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: narration, storytelling; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.6/10; 3.7s, FR.
FR_OjAMb4folEI_W000020 · in -15.4 dBFS · gain -4.6 dB · emolia-02893
(pride, disappointment· measured, subdued, steady, didactic)ce processus était excessive et donc n'était pas économiquement rentable. Aujourd'hui, on a réussi à trouver des matériaux, notamment avec des diamants artificiels,
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as pride, disappointment; style: didactic, monologue; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 0.0/10; 12.8s, FR.
FR_OjAMb4folEI_W000021 · in -17.7 dBFS · gain -2.3 dB · emolia-02893
(contempt, bitterness, anger· measured, subdued, fairly steady, cartoonish)qui font que ça devient économiquement rentable. Alors, ce qui est intéressant avec le produit, c'est que on peut désinfecter avec de l'eau, qui ressemblera pratiquement à de l'eau de javel lorsqu'elle est en phase instable, qui va faire ce travail de désinfection,
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, fairly guarded; reads as contempt, bitterness, anger; style: cartoonish, didactic; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 0.7/10; 26.9s, FR.
FR_OjAMb4folEI_W000022 · in -17.6 dBFS · gain -2.4 dB · emolia-02893
(shame·normal-paced, normally alert, fairly steady, conversational)Mais ce qui est magique, si j'ose dire, c'est qu'une fois le travail de désinfection réalisé,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame; style: conversational, storytelling; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 1.6/10; 4.1s, FR.
FR_OjAMb4folEI_W000023 · in -19.0 dBFS · gain -1.0 dB · emolia-02893
Intoxication Altered States of Consciousness ↑ (unconstrained axis: Thankfulness Gratitude)identity +0.25emotion 61 % emotion__B1__T0.60__C0.25__INTERNAL · #9
This chain comes from the one-sided rule: only Intoxication Altered States of Consciousness had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Intoxication Altered States of Consciousness below average — 0.34, lower than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.65.
Nothing was asked of the other axis, and in fact Thankfulness Gratitude drifts down from 0.95 to 0.70 (-0.25), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.15, then +0.02, then +0.24, then +0.25 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.08 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.08 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 46 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.168 before conversion and 0.414 after — it rose by 0.245. Neighbour-to-neighbour the worst pair went 0.168 → 0.414. (The earlier render, with segment 1 left raw, scores 0.432 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.651 in the original and +0.398 after conversion — 61 % of the delta retained. On the other named axis, Thankfulness Gratitude, -0.251 became -0.185.
Quality. Mean predicted overall quality across the segments went 2.39 → 2.93 (+0.54) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.168 → 0.414+0.245identity cos neighbours 0.168 → 0.414d_b rescored +0.651 → +0.398d_a rescored -0.251 → -0.185d_a mined -0.251d_b mined 0.651min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_6jEQRIzgmZMtotal 44.6schain gain +3.8 dBseam step 1.0 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice
(thankfulness gratitude, relief, embarrassment · normal-paced, energised, neutral tension, conversational)I'm sorry, Senator. Order. Order. Thank you, Senator Sheldon. Second supplementary.
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, slightly bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as thankfulness gratitude, relief, embarrassment; style: conversational, casual; below-average recording, quiet background; genuineness 4.2/6; vocal-burst blend 2.3/10; 5.7s, EN.
EN_6jEQRIzgmZM_W000343 · in -19.3 dBFS · gain -0.7 dB · emolia-01456
(measured, normally alert, slightly relaxed, monologue)Can the Minister outline the importance of having a clear plan that will grow the economy and how that will help Australians through these challenging times?
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, quiet background; genuineness 0.3/6; vocal-burst blend 0.0/10; 8.5s, EN.
EN_6jEQRIzgmZM_W000344 · in -19.5 dBFS · gain -0.5 dB · emolia-01456
(embarrassment, shame·brisk, energised, neutral tension, playful)And unlike, uh, (ahem) the previous government, which didn't have an economic plan, it just had a Prime Minister that wanted to grab up any portfolio he could.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, slightly guarded; reads as embarrassment, shame; style: playful, ranting; below-average recording, some background noise; genuineness 3.4/6; vocal-burst blend 5.4/10; 9.1s, EN.
EN_6jEQRIzgmZM_W000346 · in -18.2 dBFS · gain -1.8 dB · emolia-01456
(sourness, teasing, amusement· brisk, energised, neutral tension, cartoonish)That's the only plan that you guys had. You didn't know what each other was doing. Our economic plan is a deliberate and direct response to the economic circumstances that were left by those opposite.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, slightly rough, thin; very clear, almost no disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, fairly guarded; reads as sourness, teasing, amusement; style: cartoonish, dramatic; below-average recording, noisy background; mildly explicit content; genuineness 2.5/6; vocal-burst blend 4.3/10; 15.4s, EN.
EN_6jEQRIzgmZM_W000347 · in -17.6 dBFS · gain -2.4 dB · emolia-01456
(intoxication altered states of consciousness, teasing, impatience and irritability·normal-paced, highly aroused, tense, casual)You're hearing about it, don't you, because you know it's true. Order. Order.
full caption & clip details
A child feminine voice; delivery is highly aroused, normal-paced, tense, volatile; timbre is slightly cool, slightly bright, very rough, thin; slurred, some disfluency, very wide pitch range, heavy breath; affect is elated, slightly dominant, guarded; reads as intoxication altered states of consciousness, teasing, impatience and irritability; style: casual, dramatic; below-average recording, noisy background; mildly explicit content; genuineness 4.7/6; vocal-burst blend 2.2/10; 6.7s, EN.
EN_6jEQRIzgmZM_W000348 · in -18.2 dBFS · gain -1.8 dB · emolia-01456
This chain comes from the one-sided rule: only Infatuation had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Infatuation barely there — 0.19, lower than 81 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.67.
Nothing was asked of the other axis, and in fact Sexual Lust drifts down from 0.91 to 0.13 (-0.78), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.23, then +0.20 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.86 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.86 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 42 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.781 before conversion and 0.787 after — it rose by 0.007. Neighbour-to-neighbour the worst pair went 0.857 → 0.791. (The earlier render, with segment 1 left raw, scores 0.655 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.665 in the original and +0.096 after conversion — 14 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Sexual Lust, -0.773 became -0.211.
Quality. Mean predicted overall quality across the segments went 2.98 → 3.04 (+0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.781 → 0.787+0.007identity cos neighbours 0.857 → 0.791d_b rescored +0.665 → +0.096d_a rescored -0.773 → -0.211d_a mined -0.781d_b mined 0.665min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00024_S09529total 40.6schain gain +1.7 dBseam step 2.1 dBcrossfades 150/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, balanced body, normally alert, slightly relaxed, some disfluency
(sexual lust · measured, fairly steady, clear, didactic)So now every rectangle has been either highlighted or darkened, and if you just get rid of the darkened rectangles, you're left with just the highlighted ones, and these are your two final predictions.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as sexual lust; style: didactic, monologue; good recording, quiet background; genuineness 1.7/6; vocal-burst blend 1.9/10; 13.0s, EN.
EN_B00024_S09529_W000018 · in -23.0 dBFS · gain +3.0 dB · emolia-00724
(concentration, sexual lust, interest·normal-paced, steady, average clarity, didactic)So this is non-max suppression, and non-max means that you're going to output your maximal probabilities, classifications, but suppress the close by ones that are non-maximal. So that ends the name non-max suppression. So let's go through the details of the algorithm.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, sexual lust, interest; style: didactic, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 1.2/10; 18.6s, EN.
EN_B00024_S09529_W000019 · in -24.5 dBFS · gain +4.5 dB · emolia-00724
(measured, fairly steady, average clarity, monologue)Although for this example, I'm going to simplify it to say that you're only doing car detection.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 1.2/10; 5.0s, EN.
EN_B00024_S09529_W000020 · in -20.9 dBFS · gain +0.9 dB · emolia-00724
(measured, fairly steady, average clarity, casual)So let me get rid of the C1, C2, C3 and pretend for this slide.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 2.6/10; 4.6s, EN.
EN_B00024_S09529_W000021 · in -22.6 dBFS · gain +2.6 dB · emolia-00724
This chain comes from the one-sided rule: only Infatuation had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Infatuation barely there — 0.19, lower than 81 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.72.
Nothing was asked of the other axis, and in fact Concentration drifts down from 0.92 to 0.64 (-0.28), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.18, then +0.17, then +0.17 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 40 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.777 before conversion and 0.776 after — it fell by 0.001. Neighbour-to-neighbour the worst pair went 0.872 → 0.854. (The earlier render, with segment 1 left raw, scores 0.728 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.741 in the original and +0.645 after conversion — 87 % of the delta retained, which is most of it. On the other named axis, Concentration, -0.282 became -0.231.
Quality. Mean predicted overall quality across the segments went 3.10 → 3.23 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.777 → 0.776-0.001identity cos neighbours 0.872 → 0.854d_b rescored +0.741 → +0.645d_a rescored -0.282 → -0.231d_a mined -0.282d_b mined 0.724min_cos_consec (site) 0.8956min_cos_anchor (site) 0.9391dataset emolialang zhspeaker ZH_B00062_S06317total 39.0schain gain +1.5 dBseam step 0.7 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, no disfluency, clear
This chain comes from the one-sided rule: only Sourness had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Sourness below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.64.
Nothing was asked of the other axis, and in fact Embarrassment barely moves at all, sitting near 0.87 throughout.
It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.21, then +0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.29 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.29 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.29, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 45 s · sv · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.345 before conversion and 0.690 after — it rose by 0.344. Neighbour-to-neighbour the worst pair went 0.411 → 0.728. (The earlier render, with segment 1 left raw, scores 0.645 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sourness moved +0.640 in the original and +0.523 after conversion — 82 % of the delta retained, which is most of it. On the other named axis, Embarrassment, -0.009 became -0.010.
Quality. Mean predicted overall quality across the segments went 3.05 → 3.25 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.345 → 0.690+0.344identity cos neighbours 0.411 → 0.728d_b rescored +0.640 → +0.523d_a rescored -0.009 → -0.010d_a mined -0.012d_b mined 0.640min_cos_consec (site) 0.2911min_cos_anchor (site) 0.2911dataset podcastlang svspeaker 30542total 44.2schain gain +1.2 dBseam step 1.9 dBcrossfades 100/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, average recording, quiet background, normally alert
(measured, slightly relaxed, fairly steady, conversational)(low mumble) roligt. Men jag få fänga med skeget (ahem) för jag tänker att du Johan håller på med det ofta bostar. Du också det.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: conversational, casual; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 2.8/10; 8.8s, SV.
30542_00066992 · in -21.3 dBFS · gain +1.3 dB · podcast-01803
(longing, fatigue exhaustion, thankfulness gratitude·fast, fully relaxed, fairly steady, casual)Ja, mitt sjuk brukar växa vilken. Men jag ska ner det ganska revält annat kort och ner ord.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, fully relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as longing, fatigue exhaustion, thankfulness gratitude; style: casual, playful; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 8.2/10; 4.0s, SV.
30542_00067968 · in -21.1 dBFS · gain +1.1 dB · podcast-01806
(bitterness, interest, contempt·brisk, neutral tension, moderately variable, casual)Inget. Näligt talet, en arbetsgivare ska inte ha synpunkter på hur en arbetsta uttrycker sig eller vilka åsikter man har. Men vi måste lära sig en sak för alla. Ardsivare får inte ens ha synpunkter på arbetstagarnas eller uppdragmottagarnas politiska åsikter. Det är fyra lägger sin tid när det är professionella. Men vänta här nu.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as bitterness, interest, contempt; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 10.0/10; 19.8s, SV.
30542_00071168 · in -22.1 dBFS · gain +2.1 dB · podcast-05358
(sourness, intoxication altered states of consciousness·measured, slightly relaxed, fairly steady, whispered)Så är det ju inte. Om du på Twitter skulle skriva nu (low mumble) by the way, holocaust never happen. Då tror jag att du inte hade suttit på i talang på fred.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, intoxication altered states of consciousness; style: whispered, monologue; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 2.3/10; 12.1s, SV.
30542_00073144 · in -21.9 dBFS · gain +1.9 dB · podcast-01817
This chain comes from the one-sided rule: only Disgust had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Disgust below average — 0.33, lower than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.66.
Nothing was asked of the other axis, and in fact Awe drifts down from 1.00 to 0.47 (-0.52), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.25, then +0.21 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.82 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.82 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 31 s · zh · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.767 before conversion and 0.735 after — it fell by 0.032. Neighbour-to-neighbour the worst pair went 0.767 → 0.735. (The earlier render, with segment 1 left raw, scores 0.627 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disgust moved +0.662 in the original and +0.345 after conversion — 52 % of the delta retained. On the other named axis, Awe, -0.524 became -0.523.
Quality. Mean predicted overall quality across the segments went 2.81 → 2.98 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.767 → 0.735-0.032identity cos neighbours 0.767 → 0.735d_b rescored +0.662 → +0.345d_a rescored -0.524 → -0.523d_a mined -0.524d_b mined 0.662min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00026_S06535total 29.6schain gain +2.5 dBseam step 0.9 dBcrossfades 150/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · good recording, no background noise, clear
(infatuation, awe, pride · slow, very low-energy, relaxed, narration)Are my beautiful southern metropolis lying under the polished sky of the caribbian. No matter what say the various maps we are seen sweter even then in the island.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is warm, dark, rough, very full; clear, almost no disfluency, fairly narrow pitch, normal breath; affect is mildly negative, dominant, fairly guarded; reads as infatuation, awe, pride; style: narration, storytelling; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.2/10; 11.7s, ZH.
ZH_B00026_S06535_W000000 · in -19.6 dBFS · gain -0.5 dB · emolia-03537
(pleasure ecstasy, sexual lust, awe · slow, very low-energy, slightly relaxed, whispered)Sweeping gently over the inevitable crowds of ocean ry.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is warm, dark, very rough, very full; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, dominant, neutral openness; reads as pleasure ecstasy, sexual lust, awe; style: whispered, storytelling; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 4.1/10; 3.6s, ZH.
ZH_B00026_S06535_W000001 · in -19.6 dBFS · gain -0.4 dB · emolia-03537
(emotional numbness, longing, sourness·measured, normally alert, slightly relaxed, narration)Hurrying through the fancy art deco lovby of the park, central and to the rooms. I kept there.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, longing, sourness; style: narration, formal; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.6/10; 4.9s, ZH.
ZH_B00026_S06535_W000002 · in -18.8 dBFS · gain -1.2 dB · emolia-03537
(disgust, emotional numbness, embarrassment· measured, very low-energy, slightly relaxed, narration)I stripped off my jungle, warn clothes, and when into my own closets for a white turtle, next shirt, belted, kcke jacket and pants, and a pair of smooth brown leather boots.
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, slightly relaxed, steady; timbre is slightly warm, slightly dark, fairly smooth, very full; clear, almost no disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, neutral openness; reads as disgust, emotional numbness, embarrassment; style: narration, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 1.3/10; 10.0s, ZH.
ZH_B00026_S06535_W000003 · in -20.6 dBFS · gain +0.6 dB · emolia-03537
This chain comes from the one-sided rule: only Impatience and Irritability had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Impatience and Irritability below average — 0.38, lower than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.61.
Nothing was asked of the other axis, and in fact Fatigue Exhaustion drifts down from 0.95 to 0.73 (-0.22), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are -0.03, then +0.25, then +0.20, then +0.18 — not a clean run: step 1 moves back the other way by 0.03 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 45 s · fr · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.744 before conversion and 0.814 after — it rose by 0.070. Neighbour-to-neighbour the worst pair went 0.797 → 0.846. (The earlier render, with segment 1 left raw, scores 0.665 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.609 in the original and +0.446 after conversion — 73 % of the delta retained, which is most of it. On the other named axis, Fatigue Exhaustion, -0.218 became -0.124.
Quality. Mean predicted overall quality across the segments went 2.97 → 3.11 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.744 → 0.814+0.070identity cos neighbours 0.797 → 0.846d_b rescored +0.609 → +0.446d_a rescored -0.218 → -0.124d_a mined -0.218d_b mined 0.609min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang frspeaker FR_GHv_U5pKopktotal 43.5schain gain +1.3 dBseam step 2.3 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: an elderly masculine voice · average recording
(fatigue exhaustion, fear, disgust · measured, very low-energy, neutral tension, storytelling)Eux aussi dans la transition, (low mumble) euh, climatique, énergétique, euh, de société. Et, et à ce moment-là, de, d'avoir un, un, un débat commun sur nos désirs profonds.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, very wide pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, fear, disgust; style: storytelling, conversational; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 3.0/10; 12.4s, FR.
FR_GHv_U5pKopk_W000043 · in -20.1 dBFS · gain +0.1 dB · emolia-02891
(measured, normally alert, slightly relaxed, didactic)Le compte, (ahem) hein, euh, (low mumble) qui marche très bien pour les monnaies locales complémentaires ou les selles, c'est le compte, euh, (low mumble) d'une femme qui vient d'ailleurs.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, conversational; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 4.7/10; 8.1s, FR.
FR_GHv_U5pKopk_W000044 · in -18.6 dBFS · gain -1.4 dB · emolia-02891
(awe, astonishment surprise, longing·fast, energised, neutral tension, storytelling)qui reconnaît son village trente ans après, et qui dit c'est incroyable, la mairie c'est toujours la même, l'église c'est toujours la même, et
full caption & clip details
An elderly masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, audible breath; affect is mildly negative, slightly dominant, neutral openness; reads as awe, astonishment surprise, longing; style: storytelling, dramatic; average recording, no background noise; genuineness 2.0/6; vocal-burst blend 7.9/10; 8.3s, FR.
FR_GHv_U5pKopk_W000045 · in -18.0 dBFS · gain -2.0 dB · emolia-02891
(relief, disgust, distress·measured, energised, neutral tension, storytelling)Par contre, l'auberge, ça y était pas. Elle va voir monsieur l'aubergiste. Ça fait 30 ans que je suis pas venu dans mon vieux village. Je voudrais la plus belle chambre que vous avez. Oh, madame!
full caption & clip details
An elderly masculine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is warm, neutral-bright, slightly rough, very thin; average clarity, some disfluency, wide pitch range, audible breath; affect is positive, slightly dominant, slightly guarded; reads as relief, disgust, distress; style: storytelling, narration; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 7.3/10; 11.7s, FR.
FR_GHv_U5pKopk_W000046 · in -18.8 dBFS · gain -1.2 dB · emolia-02891
(impatience and irritability, bitterness, anger·brisk, energised, neutral tension, storytelling)pas de problème, lit la dame. Elle sort un billet de 100 euros, elle le met sur la table.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, rough, thin; average clarity, no disfluency, wide pitch range, light breath; affect is negative, slightly dominant, guarded; reads as impatience and irritability, bitterness, anger; style: storytelling, dramatic; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 6.0/10; 3.8s, FR.
FR_GHv_U5pKopk_W000047 · in -17.3 dBFS · gain -2.7 dB · emolia-02891
Intoxication Altered States of Consciousness ↑ (unconstrained axis: Fear)identity −0.21emotion 107 % emotion__B1__T0.60__C0.25__INTERNAL · #15
This chain comes from the one-sided rule: only Intoxication Altered States of Consciousness had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Intoxication Altered States of Consciousness essentially absent — 0.04, lower than 96 % of clips in this corpus — and ends with it strongly present at 0.78, higher than 78 % of clips in this corpus. That is a total rise of 0.74.
Nothing was asked of the other axis, and in fact Fear drifts down from 0.97 to 0.56 (-0.41), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.09, then +0.23, then +0.24, then +0.18 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.71 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.71 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 57 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.721 before conversion and 0.511 after — it fell by 0.210. Neighbour-to-neighbour the worst pair went 0.684 → 0.595. (The earlier render, with segment 1 left raw, scores 0.427 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.735 in the original and +0.785 after conversion — 107 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Fear, -0.413 became -0.438.
Quality. Mean predicted overall quality across the segments went 2.98 → 3.02 (+0.04) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.721 → 0.511-0.210identity cos neighbours 0.684 → 0.595d_b rescored +0.735 → +0.785d_a rescored -0.413 → -0.438d_a mined -0.413d_b mined 0.737min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00021_S01127total 55.4schain gain +2.5 dBseam step 3.0 dBcrossfades 150/100/100/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, good recording, clear, moderate pitch range, light breath
(fear, emotional numbness, astonishment surprise · measured, normally alert, slightly relaxed, narration)He attached large heavy objects to his handcuffs, so he'd pupuled down to the bottom of the water faster or sometimes after being tied up, he'd also been nailed inside a wooden box, which was then dropped in the river.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, emotional numbness, astonishment surprise; style: narration, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 14.8s, ZH.
ZH_B00021_S01127_W000000 · in -19.1 dBFS · gain -0.9 dB · emolia-03489
(awe, jealousy and envy, emotional numbness · measured, normally alert, slightly relaxed, narration)He did bridge jumps in cities all over the US from rochester, new york to boston massacsetts to neullan's louisiana. There was no way to charge money for these kinds of stuts. But harry didn't mind, as with his earlier police challenges, he was excited by the fame and attention. It wasn't hard to advertise himself now, tens of thousands of people would come out to see one of his jumps.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as awe, jealousy and envy, emotional numbness; style: narration, newsreading; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 1.1/10; 27.7s, ZH.
ZH_B00021_S01127_W000001 · in -19.7 dBFS · gain -0.3 dB · emolia-03489
(pride, awe, triumph· measured, normally alert, slightly relaxed, monologue)There was nothing who deny likes better than a big crowd.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as pride, awe, triumph; style: monologue, narration; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 5.3/10; 4.0s, ZH.
ZH_B00021_S01127_W000002 · in -20.4 dBFS · gain +0.3 dB · emolia-03489
(affection, jealousy and envy, infatuation·slow, very low-energy, relaxed, monologue)Harry loved being famous, he likes being the best. He also liked being first.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as affection, jealousy and envy, infatuation; style: monologue, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 3.0/10; 6.6s, ZH.
ZH_B00021_S01127_W000003 · in -18.5 dBFS · gain -1.5 dB · emolia-03489
(normal-paced, normally alert, slightly relaxed, formal)Airplanes came into use during harris lifetime.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: formal, casual; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 2.0/10; 3.1s, ZH.
ZH_B00021_S01127_W000004 · in -17.8 dBFS · gain -2.2 dB · emolia-03489
This chain comes from the one-sided rule: only Infatuation had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Infatuation barely there — 0.20, lower than 80 % of clips in this corpus — and ends with it strongly present at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.62.
Nothing was asked of the other axis, and in fact Emotional Numbness drifts down from 0.88 to 0.75 (-0.12), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.25, then +0.09, then +0.04 — a plateau around step 4, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.67 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.67 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 36 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.846 before conversion and 0.781 after — it fell by 0.064. Neighbour-to-neighbour the worst pair went 0.856 → 0.795. (The earlier render, with segment 1 left raw, scores 0.642 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.619 in the original and +0.559 after conversion — 90 % of the delta retained, which is essentially all of it. On the other named axis, Emotional Numbness, -0.123 became -0.247.
Quality. Mean predicted overall quality across the segments went 2.91 → 3.00 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.846 → 0.781-0.064identity cos neighbours 0.856 → 0.795d_b rescored +0.619 → +0.559d_a rescored -0.123 → -0.247d_a mined -0.123d_b mined 0.619min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_T-agynfuRV0total 34.4schain gain +2.6 dBseam step 0.8 dBcrossfades 100/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(no disfluency, formal, authoritative)In some cases, the calculations are simple, in others they are extremely complicated
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.3/10; 4.2s, EN.
EN_T-agynfuRV0_W000102 · in -14.6 dBFS · gain -5.4 dB · emolia-00418
(concentration·almost no disfluency, formal, newsreading)There is an alternative, simple method of finding the positions of the hour lines which can be used for many types of sundial, and saves a lot of work in cases where the calculations are complex.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: formal, newsreading; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.4/10; 8.9s, EN.
EN_T-agynfuRV0_W000103 · in -14.0 dBFS · gain -6.0 dB · emolia-00418
(concentration ·no disfluency, formal, authoritative)The equation of time must be taken into account to ensure that the positions of the hour lines are independent of the time of year when they are marked.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: formal, authoritative; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.0/10; 6.5s, EN.
EN_T-agynfuRV0_W000104 · in -13.8 dBFS · gain -6.2 dB · emolia-00418
(almost no disfluency, authoritative, formal)An easy way to do this is to set a clock or watch so it shows, sundial time, which is standard time, plus the equation of time on the day in question.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.4/10; 8.0s, EN.
EN_T-agynfuRV0_W000105 · in -14.7 dBFS · gain -5.3 dB · emolia-00418
(no disfluency, formal, newsreading)The hour lines on the sundial are marked to show the positions of the shadow of the style when this clock shows whole numbers of hours, and are labeled with these numbers of hours
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.1/10; 7.4s, EN.
EN_T-agynfuRV0_W000106 · in -13.3 dBFS · gain -6.7 dB · emolia-00418
This chain comes from the one-sided rule: only Confusion had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Confusion below average — 0.34, lower than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.61.
Nothing was asked of the other axis, and in fact Triumph drifts down from 0.92 to 0.15 (-0.77), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.15, then +0.24, then +0.22 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 37 s · zh · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.792 before conversion and 0.841 after — it rose by 0.049. Neighbour-to-neighbour the worst pair went 0.757 → 0.760. (The earlier render, with segment 1 left raw, scores 0.763 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.607 in the original and +0.087 after conversion — 14 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Triumph, -0.771 became -0.313.
Quality. Mean predicted overall quality across the segments went 3.02 → 3.22 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.792 → 0.841+0.049identity cos neighbours 0.757 → 0.760d_b rescored +0.607 → +0.087d_a rescored -0.771 → -0.313d_a mined -0.771d_b mined 0.606min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00077_S05689total 36.3schain gain +1.8 dBseam step 0.7 dBcrossfades 100/100/100 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, average recording, normally alert, some disfluency, moderate pitch range
This chain comes from the one-sided rule: only Impatience and Irritability had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Impatience and Irritability barely there — 0.24, lower than 76 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.63.
Nothing was asked of the other axis, and in fact Fatigue Exhaustion drifts down from 0.85 to 0.50 (-0.35), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.07, then +0.24, then +0.23, then +0.08 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 64 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.776 before conversion and 0.638 after — it fell by 0.138. Neighbour-to-neighbour the worst pair went 0.776 → 0.638. (The earlier render, with segment 1 left raw, scores 0.601 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.630 in the original and +0.377 after conversion — 60 % of the delta retained. On the other named axis, Fatigue Exhaustion, -0.352 became +0.364.
Quality. Mean predicted overall quality across the segments went 2.78 → 3.02 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.776 → 0.638-0.138identity cos neighbours 0.776 → 0.638d_b rescored +0.630 → +0.377d_a rescored -0.352 → +0.364d_a mined -0.352d_b mined 0.630min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_ZV6eKCoYmCYtotal 63.0schain gain +2.0 dBseam step 3.6 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · quiet background, measured
(normally alert, slightly relaxed, fairly steady, casual)And that thickness is controlled by how much energy you had hydrogen atoms.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 1.8/10; 3.9s, EN.
EN_ZV6eKCoYmCY_W000142 · in -16.8 dBFS · gain -3.2 dB · emolia-02625
(jealousy and envy·subdued, neutral tension, fairly steady, casual)So, you can precisely control the thickness here. You can precisely control this thickness. So, you put it upside down, you have got the SOI layer on top. You know, you have the wafer like that. You put it like that, you get the, you put it like that, you get the SOI layer. Okay? Reverse it. So, now you can see, this layer thickness was 300 microns. Out of that, let us say, half a micron is gone onto this.
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is slightly cool, dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy; style: casual, monologue; below-average recording, quiet background; genuineness 4.9/6; vocal-burst blend 5.0/10; 25.4s, EN.
EN_ZV6eKCoYmCY_W000143 · in -15.6 dBFS · gain -4.4 dB · emolia-02625
(subdued, slightly relaxed, steady, didactic)So still you have got 299.5 microns of silicon, which is separated. So once this is separated out, you have got the silicon on insulator already. You have to put it upside down, of course.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 0.3/10; 13.0s, EN.
EN_ZV6eKCoYmCY_W000144 · in -18.7 dBFS · gain -1.3 dB · emolia-02625
(normally alert, slightly relaxed, fairly steady, didactic)And this wafer which was there, it is available for you, for reusing. All that you have to do is, you may have to slightly polish it. Okay. Polish it, reuse it. Start all over again, oxidize and go through everything. So, in effect, what has happened is,
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 2.4/10; 13.7s, EN.
EN_ZV6eKCoYmCY_W000145 · in -14.0 dBFS · gain -6.0 dB · emolia-02625
(normally alert, slightly relaxed, steady, didactic)You have used two wafers to start with, but one wafer is available for you to use. Ultimately, only one wafer is used.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.9/10; 7.8s, EN.
EN_ZV6eKCoYmCY_W000146 · in -14.3 dBFS · gain -5.7 dB · emolia-02625
This chain comes from the one-sided rule: only Contemplation had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Contemplation below average — 0.28, lower than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.70.
Nothing was asked of the other axis, and in fact Jealousy and Envy drifts down from 1.00 to 0.91 (-0.09), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.17, then +0.16, then +0.16 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.31 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.31 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.31, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 61 s · de · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.422 before conversion and 0.799 after — it rose by 0.377. Neighbour-to-neighbour the worst pair went 0.438 → 0.760. (The earlier render, with segment 1 left raw, scores 0.715 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.703 in the original and +0.633 after conversion — 90 % of the delta retained, which is essentially all of it. On the other named axis, Jealousy and Envy, -0.091 became -0.089.
Quality. Mean predicted overall quality across the segments went 2.90 → 3.19 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.422 → 0.799+0.377identity cos neighbours 0.438 → 0.760d_b rescored +0.703 → +0.633d_a rescored -0.091 → -0.089d_a mined -0.090d_b mined 0.696min_cos_consec (site) 0.3051min_cos_anchor (site) 0.3096dataset podcastlang despeaker 491220total 60.1schain gain +1.6 dBseam step 2.9 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, normally alert
(jealousy and envy, impatience and irritability, elation · normal-paced, neutral tension, moderately variable, conversational)hat nicht genau eine schöne Pixel-Grafik. PlayStation Plus rasiert auch wieder alles dieses Mal. EA Sports U of C4, wo wollte das nicht schon mal spielen. Planet Coaster Console Edition, ja, warum denn nicht? Tiny
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as jealousy and envy, impatience and irritability, elation; style: conversational, casual; good recording, quiet background; genuineness 4.3/6; vocal-burst blend 1.8/10; 12.0s, DE.
491220_00379576 · in -18.7 dBFS · gain -1.3 dB · podcast-02152
(disgust, astonishment surprise, intoxication altered states of consciousness· normal-paced, neutral tension, fairly steady, casual)Tina's Assault on Dragon Keep A Wonderlands One Shot, Anteger, viel zu lange Titel, deswegen nicht spielen für die PS5. (low mumble) Ist ganz schön dürftig, durch die Bank, oder? Also ich meine gut, Game Pass, aber Playstation,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as disgust, astonishment surprise, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 1.4/10; 15.2s, DE.
491220_00380784 · in -19.6 dBFS · gain -0.4 dB · podcast-02151
(disgust, jealousy and envy, sourness·measured, slightly relaxed, fairly steady, casual)aber da haben wir schon mal drüber geredet, die, nachdem die Xbox (low mumble) Live Gold immer so im Sack hatten, so langsam schwächeln die auch ganz massiv. Was soll es?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, jealousy and envy, sourness; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 0.0/10; 10.1s, DE.
491220_00382312 · in -19.5 dBFS · gain -0.5 dB · podcast-02151
(confusion, longing, sourness · measured, slightly relaxed, fairly steady, conversational)Ist dieses Tiny Tinas-Ding, ist das nicht auch neu, zumindest mein.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as confusion, longing, sourness; style: conversational, casual; average recording, no background noise; genuineness 2.6/6; vocal-burst blend 0.0/10; 4.6s, DE.
491220_00383360 · in -18.3 dBFS · gain -1.7 dB · podcast-02181
(contemplation, confusion, jealousy and envy·normal-paced, neutral tension, fairly steady, casual)Kann sein, aber es ist Borderlands, deswegen interessiert mich das nicht, die Bohne.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly submissive, slightly guarded; reads as contemplation, confusion, jealousy and envy; style: casual, conversational; average recording, quiet background; genuineness 5.8/6; vocal-burst blend 3.4/10; 18.8s, DE.
491220_00383912 · in -19.1 dBFS · gain -0.9 dB · podcast-03846
This chain comes from the one-sided rule: only Relief had to get where it was going, by at least 0.60. The other emotion was left completely free.
The chain starts with Relief below average — 0.40, lower than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.60.
Nothing was asked of the other axis, and in fact Amusement barely moves at all, sitting near 0.97 throughout.
It takes 4 clips to get there. Clip to clip the moves are +0.15, then +0.21, then +0.24 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.08 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.15 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.08, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 39 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.092 before conversion and 0.316 after — it rose by 0.225. Neighbour-to-neighbour the worst pair went 0.348 → 0.581. (The earlier render, with segment 1 left raw, scores 0.221 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.602 in the original and +0.978 after conversion — 163 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Amusement, -0.007 became -0.001.
Quality. Mean predicted overall quality across the segments went 2.78 → 3.06 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.092 → 0.316+0.225identity cos neighbours 0.348 → 0.581d_b rescored +0.602 → +0.978d_a rescored -0.007 → -0.001d_a mined -0.007d_b mined 0.602min_cos_consec (site) 0.1547min_cos_anchor (site) 0.0786dataset podcastlang enspeaker 678560total 38.3schain gain +1.4 dBseam step 1.4 dBcrossfades 100/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright
(amusement, embarrassment, confusion · normal-paced, normally alert, fully relaxed, casual)weird, but yeah, the like the question about Fallacio, I was just like,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as amusement, embarrassment, confusion; style: casual, conversational; average recording, quiet background; genuineness 5.8/6; vocal-burst blend 4.1/10; 3.8s, EN.
678560_00091992 · in -20.5 dBFS · gain +0.5 dB · podcast-05791
(sexual lust, confusion, doubt· normal-paced, normally alert, slightly relaxed, casual)wait a second, wait a second. I feel like a doofus right now. What does Falatio mean?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as sexual lust, confusion, doubt; style: casual, conversational; average recording, quiet background; genuineness 5.4/6; vocal-burst blend 2.7/10; 4.0s, EN.
678560_00093256 · in -17.2 dBFS · gain -2.8 dB · podcast-05780
(measured, normally alert, slightly relaxed, formal)this. Okay, there was one question. When was the first time in your relationship that you engaged in oral sex? Yeah,
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 0.8/10; 8.7s, EN.
678560_00094496 · in -22.4 dBFS · gain +2.4 dB · podcast-05790
(relief, thankfulness gratitude, amusement·slow, very low-energy, relaxed, conversational)Okay, excellent. I thought so, because yeah, that particular question, like we (chuckle) (chuckle) one senior (low mumble) uh individual gave the feedback that, like, oh my gosh, you know, like when you ask that question, then it just makes me think of all the times that that happened, and that's like agitating to the mind, you know. People
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as relief, thankfulness gratitude, amusement; style: conversational, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 5.2/10; 22.3s, EN.
678560_00095624 · in -23.3 dBFS · gain +3.3 dB · podcast-03196