Manifest tier. emotion, rule B1, T=0.2, step cap 0.25. Population 21,543,847 chains (214,925 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 16,505,714.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words
are a real non-speech sound, printed where it happens.
The full generated caption for any clip is under “full caption & clip
details”. Its perceived-gender and background-noise clauses were re-rendered from
the numeric buckets, because the versions stored in the corpus index had those two
ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
The tags and captions describe the ORIGINAL clips — they are the
corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the
converted audio are the ones printed in each card's own paragraph and chips.
This chain comes from the one-sided rule: only Longing had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Longing clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.28.
Nothing was asked of the other axis, and in fact Thankfulness Gratitude drifts down from 0.93 to 0.33 (-0.60), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.13, then -0.06, then +0.21 — not a clean run: step 2 moves back the other way by 0.06 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.68 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.68 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 33 s · ko · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.755 before conversion and 0.729 after — it fell by 0.026. Neighbour-to-neighbour the worst pair went 0.780 → 0.783. (The earlier render, with segment 1 left raw, scores 0.722 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.279 in the original and +0.244 after conversion — 87 % of the delta retained, which is most of it. On the other named axis, Thankfulness Gratitude, -0.601 became -0.415.
Quality. Mean predicted overall quality across the segments went 3.00 → 3.11 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.755 → 0.729-0.026identity cos neighbours 0.780 → 0.783d_b rescored +0.279 → +0.244d_a rescored -0.601 → -0.415d_a mined -0.601d_b mined 0.279min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang kospeaker KO_9nRbG8Lpa5Etotal 32.2schain gain -0.7 dBseam step 0.7 dBcrossfades 100/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range, light breath
(thankfulness gratitude, infatuation · normal-paced, no disfluency, clear, formal)너희 투자 가치가 없는 상황이 되었고 한 10년은 기다려야 안 되겠나 싶습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, infatuation; style: formal, narration; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 3.1/10; 4.3s, KO.
KO_9nRbG8Lpa5E_W000017 · in -17.4 dBFS · gain -2.5 dB · emolia-03197
(shame, fatigue exhaustion· normal-paced, some disfluency, somewhat unclear, conversational)아, 마찬가지, 남산동에 위치한 주택인데요. 확인을 해보게 되면. 그래도 뭐, 한 4m 정도는 접하고 있는 주택 같아요. 자세히 확인을 해보면.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, fatigue exhaustion; style: conversational, monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 5.8/10; 8.5s, KO.
KO_9nRbG8Lpa5E_W000018 · in -21.9 dBFS · gain +1.9 dB · emolia-03197
(normal-paced, some disfluency, somewhat unclear, monologue)실제 상세 현황 자체는 거의 나오지 않고 있고, 매매 극액 4억 8천, 대출금 1억이 있는 것으로 확인이 됩니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, authoritative; average recording, no background noise; genuineness 3.0/6; vocal-burst blend 5.4/10; 6.3s, KO.
KO_9nRbG8Lpa5E_W000019 · in -17.9 dBFS · gain -2.1 dB · emolia-03197
(longing, contemplation, confusion·measured, some disfluency, somewhat unclear, monologue)평당 금액을 확인해보면 평당 1,600만 원 정도 치입니다. 물론 8메다 기준으로 예전에 1,500, 1,600, 수승구, 남구 할 것 없이 그렇게 거래되었는데 1,600 정도의 가치가 있는지 확인을 한번 해보겠습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing, contemplation, confusion; style: monologue, conversational; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 7.9/10; 13.7s, KO.
KO_9nRbG8Lpa5E_W000020 · in -18.9 dBFS · gain -1.1 dB · emolia-03197
This chain comes from the one-sided rule: only Contentment had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Contentment clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.33.
Nothing was asked of the other axis, and in fact Concentration drifts down from 0.77 to 0.35 (-0.42), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.10, then +0.23 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.81 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.81 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 29 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.808 before conversion and 0.785 after — it fell by 0.023. Neighbour-to-neighbour the worst pair went 0.803 → 0.814. (The earlier render, with segment 1 left raw, scores 0.698 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.329 in the original and +0.277 after conversion — 84 % of the delta retained, which is most of it. On the other named axis, Concentration, -0.419 became -0.345.
Quality. Mean predicted overall quality across the segments went 2.87 → 3.07 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.808 → 0.785-0.023identity cos neighbours 0.803 → 0.814d_b rescored +0.329 → +0.277d_a rescored -0.419 → -0.345d_a mined -0.418d_b mined 0.329min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_LivCQRvY9AYtotal 28.4schain gain +5.0 dBseam step 1.2 dBcrossfades 150/100 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, normally alert, fairly steady, moderate pitch range, light breath
(normal-paced, slightly relaxed, some disfluency, conversational)So I just answered some questions there about how the payment feature works and it does run through the teacher.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: conversational, casual; good recording, no background noise; genuineness 3.0/6; vocal-burst blend 1.3/10; 6.1s, EN.
EN_LivCQRvY9AY_W000379 · in -19.7 dBFS · gain -0.3 dB · emolia-01179
(shame, pain, contemplation·measured, relaxed, frequent disfluency, casual)(low mumble) Uhm, only, and again, (ahem) at the beginning of the year I, I added it into an admin fee, (ahem) uhm,
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as shame, pain, contemplation; style: casual, monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 2.1/10; 8.1s, EN.
EN_LivCQRvY9AY_W000380 · in -17.9 dBFS · gain -2.1 dB · emolia-01179
(contentment, hope enthusiasm optimism, affection·normal-paced, slightly relaxed, some disfluency, casual)(ahem) I would suggest that parents might be open to the fact that you're going to now have an online teaching extra fee that you would implement so that you can invest in the tools you need to give their children the best opportunity that you can give them.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contentment, hope enthusiasm optimism, affection; style: casual, conversational; good recording, quiet background; genuineness 1.7/6; vocal-burst blend 1.2/10; 14.6s, EN.
EN_LivCQRvY9AY_W000381 · in -21.6 dBFS · gain +1.6 dB · emolia-01179
This chain comes from the one-sided rule: only Fear had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Fear below average — 0.27, lower than 73 % of clips in this corpus — and ends with it strongly present at 0.81, higher than 81 % of clips in this corpus. That is a total rise of 0.54.
Nothing was asked of the other axis, and in fact Concentration drifts down from 1.00 to 0.74 (-0.26), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.15, then +0.14 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.82 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.82 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 29 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.638 before conversion and 0.643 after — it rose by 0.006. Neighbour-to-neighbour the worst pair went 0.828 → 0.783. (The earlier render, with segment 1 left raw, scores 0.414 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.538 in the original and +0.531 after conversion — 99 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.262 became -0.306.
Quality. Mean predicted overall quality across the segments went 2.54 → 2.75 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.638 → 0.643+0.006identity cos neighbours 0.828 → 0.783d_b rescored +0.538 → +0.531d_a rescored -0.262 → -0.306d_a mined -0.262d_b mined 0.538min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_Wjk8SUA0tUItotal 28.3schain gain +2.4 dBseam step 0.7 dBcrossfades 100/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(concentration, triumph · somewhat unclear, casual, monologue)We are even improving the temperature forecast by more than three to four Kelvin or three to four degrees (ahem) Celsius every day. So we feel that we are getting the highest improvement on the second day of the forecast because we refresh the meteorology.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, triumph; style: casual, monologue; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.3/10; 15.5s, EN.
EN_Wjk8SUA0tUI_W000050 · in -19.4 dBFS · gain -0.6 dB · emolia-00456
(somewhat unclear, casual, conversational)Every day when we launch a new forecast. And so the older.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 2.8/10; 5.2s, EN.
EN_Wjk8SUA0tUI_W000051 · in -18.4 dBFS · gain -1.6 dB · emolia-00456
(average clarity, casual, playful)The advantage that we have accumulated over the
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, playful; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 4.1/10; 3.3s, EN.
EN_Wjk8SUA0tUI_W000052 · in -19.3 dBFS · gain -0.7 dB · emolia-00456
(average clarity, casual, monologue)Downward reaching (ahem) solar radiation at the surface and planetary boundary layer height.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 3.3/10; 4.9s, EN.
EN_Wjk8SUA0tUI_W000053 · in -19.7 dBFS · gain -0.3 dB · emolia-00456
This chain comes from the one-sided rule: only Astonishment Surprise had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Astonishment Surprise strongly present — 0.78, higher than 78 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.21.
Nothing was asked of the other axis, and in fact Thankfulness Gratitude drifts down from 0.99 to 0.62 (-0.37), which the rule did not require.
It takes 2 clips to get there. Clip to clip the moves are +0.21 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 12 s · zh · emolia
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.531 before conversion and 0.691 after — it rose by 0.160. Neighbour-to-neighbour the worst pair went 0.531 → 0.691. (The earlier render, with segment 1 left raw, scores 0.680 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.211 in the original and +0.234 after conversion — 111 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Thankfulness Gratitude, -0.371 became -0.150.
Quality. Mean predicted overall quality across the segments went 2.95 → 3.00 (+0.04) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.531 → 0.691+0.160identity cos neighbours 0.531 → 0.691d_b rescored +0.211 → +0.234d_a rescored -0.371 → -0.150d_a mined -0.371d_b mined 0.210min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00066_S09603total 11.5schain gain +0.5 dBseam step 0.0 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-bright, good recording, no background noise, measured, normally alert, slightly relaxed, clear, wide pitch range
(thankfulness gratitude, jealousy and envy, teasing · moderately variable, little disfluency, playful, storytelling)I hope he won't do that without your permission, said mrs. Penny man.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as thankfulness gratitude, jealousy and envy, teasing; style: playful, storytelling; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.3/10; 6.1s, ZH.
ZH_B00066_S09603_W000192 · in -21.6 dBFS · gain +1.6 dB · emolia-03935
(astonishment surprise, affection, longing·fairly steady, almost no disfluency, storytelling, whispered)My dear, he seems to have yours her brother answered.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, thin; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as astonishment surprise, affection, longing; style: storytelling, whispered; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.9/10; 5.6s, ZH.
ZH_B00066_S09603_W000193 · in -20.1 dBFS · gain +0.1 dB · emolia-03935
This chain comes from the one-sided rule: only Affection had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Affection strongly present — 0.77, higher than 77 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.22.
Nothing was asked of the other axis, and in fact Pride drifts down from 1.00 to 0.94 (-0.06), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.00, then +0.22 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.40 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.41 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.40, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 48 s · pt · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.389 before conversion and 0.767 after — it rose by 0.378. Neighbour-to-neighbour the worst pair went 0.396 → 0.804. (The earlier render, with segment 1 left raw, scores 0.673 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Affection moved +0.226 in the original and +0.353 after conversion — 156 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Pride, -0.061 became -0.063.
Quality. Mean predicted overall quality across the segments went 2.53 → 3.33 (+0.80) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.389 → 0.767+0.378identity cos neighbours 0.396 → 0.804d_b rescored +0.226 → +0.353d_a rescored -0.061 → -0.063d_a mined -0.058d_b mined 0.225min_cos_consec (site) 0.4078min_cos_anchor (site) 0.4011dataset podcastlang ptspeaker 469559total 46.9schain gain +1.9 dBseam step 1.4 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · relaxed, frequent disfluency, audible breath
(pride, contentment, doubt · measured, very low-energy, fairly steady, monologue)E mesmo assim me prende, me dá vontade de ficar olhando. Não sei se é pelo costume do YouTube ou algo do tipo. Mas se eu estiver no Spotify, eu sei que eu não preciso ficar olhando. Mas no YouTube não. Aí eu sinto que existe uma mínima necessidade de ter algo, entende? Pra visualizar.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as pride, contentment, doubt; style: monologue, whispered; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 4.2/10; 19.6s, PT.
469559_00051755 · in -32.0 dBFS · gain +12.1 dB · podcast-05519
(embarrassment, infatuation, astonishment surprise·normal-paced, normally alert, moderately variable, casual)Não, mas se liga. Estou fazendo feijão na panela de pressão. Ah, eu vi os stories, velho. Boy, eu nunca imaginei que eu fosse conseguir fazer isso. É claro que tipo, todo mundo aqui me ajuda, tá ligado? É quando
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, very dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, audible breath; affect is positive, slightly submissive, neutral openness; reads as embarrassment, infatuation, astonishment surprise; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 6.0/6; vocal-burst blend 9.0/10; 13.4s, PT.
469559_00076936 · in -27.2 dBFS · gain +7.2 dB · podcast-04937
This chain comes from the one-sided rule: only Thankfulness Gratitude had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Thankfulness Gratitude strongly present — 0.77, higher than 77 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.23.
Nothing was asked of the other axis, and in fact Pride barely moves at all, sitting near 0.99 throughout.
It takes 4 clips to get there. Clip to clip the moves are +0.08, then +0.11, then +0.04 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 63 s · hr · eurospeech
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.885 before conversion and 0.909 after — it rose by 0.024. Neighbour-to-neighbour the worst pair went 0.884 → 0.871. (The earlier render, with segment 1 left raw, scores 0.722 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.224 in the original and +0.968 after conversion — 432 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Pride, -0.019 became -0.259.
Quality. Mean predicted overall quality across the segments went 2.97 → 3.21 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.885 → 0.909+0.024identity cos neighbours 0.884 → 0.871d_b rescored +0.224 → +0.968d_a rescored -0.019 → -0.259d_a mined -0.019d_b mined 0.226min_cos_consec (site) —min_cos_anchor (site) —dataset eurospeechlang hrspeaker croatia_20081120161001-273total 62.1schain gain +1.6 dBseam step 0.2 dBcrossfades 100/100/100 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · slightly cool, fairly smooth, thin, average recording, neutral tension, moderately variable, some disfluency, light breath
(pride, triumph, interest · brisk, normally alert, average clarity, cartoonish)je svakako tehnička pomoć koja pomaže upravljanju ovim operativnim programom što isto nije nebitno dakle osim što je bitno da imamo dobre projekte bitno je u svakom slučaju i da ih dobro (ahem) provodimo. Evo za kraj mislim da je ovo jedan projekt, jedan dio programa IPA
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pride, triumph, interest; style: cartoonish, ranting; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 6.7/10; 17.5s, HR.
croatia_20081120161001-2735_1257168_1274656 · in -26.1 dBFS · gain +6.1 dB · eurospeech-01261
(shame, triumph, hope enthusiasm optimism· brisk, normally alert, average clarity, ranting)koji radi u isto vrijeme i na koheziji i na konkurentnosti. Dakle priprema Hrvatsku i po onim raznim uvjetima za članstvo. Ja bih ovdje osobito istaknula onaj, onaj drugi uvjet a to je (ahem) ekonomski dakle koji govori o mogućnosti dakle tržišnom gospodarstvu koje eto jesmo već,
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as shame, triumph, hope enthusiasm optimism; style: ranting, cartoonish; average recording, some background noise; genuineness 3.1/6; vocal-burst blend 4.2/10; 19.3s, HR.
croatia_20081120161001-2735_1274656_1293936 · in -25.5 dBFS · gain +5.5 dB · eurospeech-01261
(thankfulness gratitude, infatuation· brisk, energised, clear, cartoonish)ali u onom drugom dijelu o kome možda se dovoljno često ne govori o razvoju konkurentnosti. Dakle Hrvatska mora biti spremna nositi se sa izazovima koje će donijeti otvaranje tržišta
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, guarded; reads as thankfulness gratitude, infatuation; style: cartoonish, ranting; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 6.2/10; 11.0s, HR.
croatia_20081120161001-2735_1293936_1304928 · in -25.3 dBFS · gain +5.3 dB · eurospeech-01261
(thankfulness gratitude, triumph, pride·normal-paced, normally alert, average clarity, casual)i mislim da ovaj projekt koliko god nije možda najveći od svih ipak komponenti čini i jedan bitan dio jer će se odnositi na područja koja sigurno trebaju najveću pomoć u Republici Hrvatskoj . Hvala lijepa.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as thankfulness gratitude, triumph, pride; style: casual, monologue; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 2.9/10; 14.8s, HR.
croatia_20081120161001-2735_1304928_1319696 · in -27.2 dBFS · gain +7.2 dB · eurospeech-01261
This chain comes from the one-sided rule: only Pleasure Ecstasy had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Pleasure Ecstasy clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.24.
Nothing was asked of the other axis, and in fact Embarrassment barely moves at all, sitting near 0.99 throughout.
It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.02 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.71 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.71 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.71, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 12 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.704 before conversion and 0.495 after — it fell by 0.209. Neighbour-to-neighbour the worst pair went 0.704 → 0.485. (The earlier render, with segment 1 left raw, scores 0.287 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pleasure Ecstasy moved +0.244 in the original and +0.413 after conversion — 170 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Embarrassment, -0.022 became -0.037.
Quality. Mean predicted overall quality across the segments went 2.45 → 2.69 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.704 → 0.495-0.209identity cos neighbours 0.704 → 0.485d_b rescored +0.244 → +0.413d_a rescored -0.022 → -0.037d_a mined -0.022d_b mined 0.239min_cos_consec (site) 0.7061min_cos_anchor (site) 0.7061dataset podcastlang enspeaker 846343total 11.5schain gain +1.7 dBseam step 1.5 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, slightly bright, fairly smooth, balanced body, normal-paced, normally alert, moderately variable, some disfluency
(embarrassment, astonishment surprise, sexual lust · slightly relaxed, conversational, casual)but it came off like a little bit too much. (wistful sigh) Yeah. Um cocky.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as embarrassment, astonishment surprise, sexual lust; style: conversational, casual; average recording, quiet background; mildly explicit content; genuineness 4.2/6; vocal-burst blend 2.4/10; 3.8s, EN.
846343_00112896 · in -26.6 dBFS · gain +6.6 dB · podcast-01073
(infatuation, amusement, contentment·relaxed, storytelling, conversational)One of my favorites though was Easy. I think he's hilarious.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as infatuation, amusement, contentment; style: storytelling, conversational; good recording, no background noise; genuineness 2.8/6; vocal-burst blend 4.0/10; 3.8s, EN.
846343_00113424 · in -21.5 dBFS · gain +1.5 dB · podcast-00161
(pleasure ecstasy, sexual lust, amusement ·neutral tension, casual, conversational)kind of that just like super funny, like almost class clown type that we
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as pleasure ecstasy, sexual lust, amusement; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 3.8/6; vocal-burst blend 8.5/10; 4.3s, EN.
846343_00113952 · in -21.5 dBFS · gain +1.5 dB · podcast-01074
This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Concentration clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.27.
Nothing was asked of the other axis, and in fact Emotional Numbness barely moves at all, sitting near 0.75 throughout.
It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.10 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.91 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.91 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 30 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.891 before conversion and 0.833 after — it fell by 0.058. Neighbour-to-neighbour the worst pair went 0.937 → 0.892. (The earlier render, with segment 1 left raw, scores 0.677 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.273 in the original and +0.288 after conversion — 105 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, -0.015 became +0.086.
Quality. Mean predicted overall quality across the segments went 2.97 → 3.12 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.891 → 0.833-0.058identity cos neighbours 0.937 → 0.892d_b rescored +0.273 → +0.288d_a rescored -0.015 → +0.086d_a mined -0.015d_b mined 0.274min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_gibgPKdzOaototal 29.7schain gain +0.2 dBseam step 0.4 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fairly steady, no disfluency, formal, monologue)Another major area of neuroscience is directed at investigations of the development of the nervous system
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.5/10; 5.4s, EN.
EN_gibgPKdzOao_W000063 · in -14.6 dBFS · gain -5.4 dB · emolia-01580
(fairly steady, no disfluency, newsreading, formal)Computational neurogenetic modeling is concerned with the development of dynamic neuronal models for modeling brain functions with respect to genes and dynamic interactions between genes
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.7/10; 9.4s, EN.
EN_gibgPKdzOao_W000065 · in -12.6 dBFS · gain -7.4 dB · emolia-01580
(concentration·steady, almost no disfluency, newsreading, authoritative)At the systems level, the questions addressed in systems neuroscience include how neural circuits are formed and used anatomically and physiologically to produce functions such as reflexes, multisensory integration, motor coordination, circadian rhythms, emotional responses, learning, and memory.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: newsreading, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 15.3s, EN.
EN_gibgPKdzOao_W000067 · in -13.1 dBFS · gain -6.9 dB · emolia-01580
This chain comes from the one-sided rule: only Longing had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Longing around average — 0.54, higher than 54 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.45.
Nothing was asked of the other axis, and in fact Contentment drifts down from 0.94 to 0.72 (-0.23), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.24 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 44 s · dutch · mls
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.860 before conversion and 0.909 after — it rose by 0.049. Neighbour-to-neighbour the worst pair went 0.881 → 0.905. (The earlier render, with segment 1 left raw, scores 0.827 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.458 in the original and +0.297 after conversion — 65 % of the delta retained. On the other named axis, Contentment, -0.225 became -0.079.
Quality. Mean predicted overall quality across the segments went 3.11 → 3.32 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.860 → 0.909+0.049identity cos neighbours 0.881 → 0.905d_b rescored +0.458 → +0.297d_a rescored -0.225 → -0.079d_a mined -0.225d_b mined 0.451min_cos_consec (site) —min_cos_anchor (site) —dataset mlslang dutchspeaker 1666total 42.8schain gain +1.0 dBseam step 0.3 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(contentment, disgust · almost no disfluency, clear, wide pitch range, narration)had haar mond niet tegen hem opengedaan de détails van den twist wist men niet goed alleen was men er zeker van dat eline s nachts in dien storm met een nachtwacht en een jong mensch in een rijtuig gezeten had en men vond dat minstens genomen vreemd
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as contentment, disgust; style: narration, storytelling; good recording, quiet background; genuineness 0.6/6; vocal-burst blend 0.1/10; 14.8s, DUTCH.
1666_1841_002325 · in -26.6 dBFS · gain +6.6 dB · mls-00099
(concentration, interest·some disfluency, average clarity, moderate pitch range, casual)enfin eline was altijd nogal excentriek geweest des winters ging ze alleen ochtendwandelingen maken in het bosch welk fatsoenlijk jong meisje deed dat nu die geschiedenis met erlevoort was ook toch nogal duister en nu die roman met een jongmensch en een nachtwacht
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, interest; style: casual, narration; good recording, quiet background; genuineness 2.6/6; vocal-burst blend 2.3/10; 14.3s, DUTCH.
1666_1841_003174 · in -23.9 dBFS · gain +3.9 dB · mls-00099
(longing, jealousy and envy, contemplation· some disfluency, average clarity, moderate pitch range, whispered)het was zoo jammer want ze was toch au fond zoo lief zoo mooi en zoo elegant maar was het niet altijd een vreemde familie geweest bij die vere's betsy verbeet zich van nijdigheid over deze praatjes waarvan zij het geruisch als in de lucht ried
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as longing, jealousy and envy, contemplation; style: whispered, casual; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 0.8/10; 14.1s, DUTCH.
1666_1841_002015 · in -27.6 dBFS · gain +7.7 dB · mls-00099
This chain comes from the one-sided rule: only Longing had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Longing strongly present — 0.76, higher than 76 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.23.
Nothing was asked of the other axis, and in fact Jealousy and Envy drifts down from 1.00 to 0.66 (-0.34), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.22, then +0.00, then +0.01 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.54 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.54 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 27 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.520 before conversion and 0.737 after — it rose by 0.217. Neighbour-to-neighbour the worst pair went 0.438 → 0.666. (The earlier render, with segment 1 left raw, scores 0.541 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.231 in the original and +0.181 after conversion — 78 % of the delta retained, which is most of it. On the other named axis, Jealousy and Envy, -0.336 became +0.015.
Quality. Mean predicted overall quality across the segments went 2.52 → 2.92 (+0.40) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.520 → 0.737+0.217identity cos neighbours 0.438 → 0.666d_b rescored +0.231 → +0.181d_a rescored -0.336 → +0.015d_a mined -0.336d_b mined 0.231min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_8NXC1PatUAQtotal 25.9schain gain +2.9 dBseam step 1.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice
(jealousy and envy, impatience and irritability, amusement · brisk, energised, neutral tension, cartoonish)I don't think women ought to sit down at table with men. (ahem) Oh, don't you? Why not? It ruins conversation.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, very rough, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is negative, slightly dominant, guarded; reads as jealousy and envy, impatience and irritability, amusement; style: cartoonish, storytelling; below-average recording, quiet background; genuineness 2.2/6; vocal-burst blend 1.5/10; 5.6s, EN.
EN_8NXC1PatUAQ_W000331 · in -16.6 dBFS · gain -3.5 dB · emolia-02606
(intoxication altered states of consciousness, longing, fatigue exhaustion·measured, highly aroused, tense, storytelling)Don't stand by my chair in order to make eyes at him. Better get Philip some more ale.
full caption & clip details
A middle-aged strongly masculine voice; delivery is highly aroused, measured, tense, moderately variable; timbre is slightly cool, dark, rough, thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is negative, dominant, guarded; reads as intoxication altered states of consciousness, longing, fatigue exhaustion; style: storytelling, cartoonish; poor recording, some background noise; genuineness 2.4/6; vocal-burst blend 0.7/10; 5.8s, EN.
EN_8NXC1PatUAQ_W000340 · in -17.7 dBFS · gain -2.3 dB · emolia-02606
(jealousy and envy, bitterness, contempt·brisk, highly aroused, tense, storytelling)Hang it all you other wife who can cook your dinner and look after your children. Don't you think so, Sally?
full caption & clip details
A middle-aged strongly masculine voice; delivery is highly aroused, brisk, tense, moderately variable; timbre is cool, neutral-bright, rough, thin; very clear, almost no disfluency, very wide pitch range, light breath; affect is negative, dominant, fairly guarded; reads as jealousy and envy, bitterness, contempt; style: storytelling, cartoonish; poor recording, quiet background; genuineness 0.7/6; vocal-burst blend 0.8/10; 5.9s, EN.
EN_8NXC1PatUAQ_W000347 · in -16.5 dBFS · gain -3.5 dB · emolia-02606
(longing, shame, distress·measured, very low-energy, neutral tension, casual)You don't know what this means to me. You see, I've practically never had any family. This is almost the only place I've ever known that's had the quality of... of home.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, neutral tension, variable; timbre is slightly cool, dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, slightly vulnerable; reads as longing, shame, distress; style: casual, conversational; below-average recording, quiet background; genuineness 3.6/6; vocal-burst blend 3.9/10; 9.2s, EN.
EN_8NXC1PatUAQ_W000351 · in -20.5 dBFS · gain +0.5 dB · emolia-02606
This chain comes from the one-sided rule: only Bitterness had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Bitterness around average — 0.56, higher than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.36.
Nothing was asked of the other axis, and in fact Fear drifts down from 0.98 to 0.35 (-0.63), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.23 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 23 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.824 before conversion and 0.815 after — it fell by 0.008. Neighbour-to-neighbour the worst pair went 0.819 → 0.793. (The earlier render, with segment 1 left raw, scores 0.768 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Bitterness moved +0.360 in the original and +0.270 after conversion — 75 % of the delta retained, which is most of it. On the other named axis, Fear, -0.628 became -0.710.
Quality. Mean predicted overall quality across the segments went 2.83 → 2.95 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.824 → 0.815-0.008identity cos neighbours 0.819 → 0.793d_b rescored +0.360 → +0.270d_a rescored -0.628 → -0.710d_a mined -0.628d_b mined 0.360min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_v9ZYAceqajMtotal 21.7schain gain -0.4 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · good recording, no background noise, steady, no disfluency, clear, fairly narrow pitch
(fear, distress, affection · slow, very low-energy, slightly relaxed, whispered)But, if you leave your job, and you prefer to hang out, you will be even more stressed after you hang out.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is slightly warm, dark, slightly rough, thin; clear, no disfluency, fairly narrow pitch, light breath; affect is mildly positive, slightly submissive, fairly guarded; reads as fear, distress, affection; style: whispered, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 1.0/10; 7.4s, EN.
EN_v9ZYAceqajM_W000022 · in -14.6 dBFS · gain -5.4 dB · emolia-02335
(fear, distress, sadness· slow, very low-energy, relaxed, whispered)But you should know that there are many of your friends who are still working hard to be able to take your position or will even surpass you.
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, smooth, very full; clear, no disfluency, fairly narrow pitch, minimal breath; affect is mildly positive, submissive, neutral openness; reads as fear, distress, sadness; style: whispered, ASMR; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.0/10; 9.5s, EN.
EN_v9ZYAceqajM_W000024 · in -15.0 dBFS · gain -5.0 dB · emolia-02335
(bitterness·measured, subdued, slightly relaxed, whispered)Discipline is the difference between people who talk big and people who work big.
full caption & clip details
A young adult feminine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, neutral openness; reads as bitterness; style: whispered, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.7/10; 5.3s, EN.
EN_v9ZYAceqajM_W000025 · in -13.3 dBFS · gain -6.7 dB · emolia-02335
This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Emotional Numbness clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.22.
Nothing was asked of the other axis, and in fact Concentration drifts down from 0.98 to 0.11 (-0.87), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.14, then +0.09 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.70 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.70 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 26 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.711 before conversion and 0.656 after — it fell by 0.055. Neighbour-to-neighbour the worst pair went 0.803 → 0.686. (The earlier render, with segment 1 left raw, scores 0.539 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.221 in the original and +0.209 after conversion — 94 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.871 became -0.831.
Quality. Mean predicted overall quality across the segments went 2.77 → 3.04 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.711 → 0.656-0.055identity cos neighbours 0.803 → 0.686d_b rescored +0.221 → +0.209d_a rescored -0.871 → -0.831d_a mined -0.871d_b mined 0.222min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_zzOFZSRqeMctotal 25.4schain gain +4.3 dBseam step 0.3 dBcrossfades 150/100 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, average clarity
(concentration, contemplation · normal-paced, some disfluency, light breath, monologue)Like for the, the, the number of, the minimum number of contributions before someone is no longer an occasional contributor, that filter doesn't differentiate between that person, whether they stopped contributing or whether they became a core contributor.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, contemplation; style: monologue, casual; good recording, no background noise; genuineness 3.0/6; vocal-burst blend 1.7/10; 13.4s, EN.
EN_zzOFZSRqeMc_W000123 · in -19.1 dBFS · gain -0.9 dB · emolia-01998
(disgust, contempt, bitterness· normal-paced, some disfluency, normal breath, casual)Right, so there's no, there's no value attributed to those filters (low mumble) as far as directionality goes.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, contempt, bitterness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 3.9/6; vocal-burst blend 1.2/10; 7.0s, EN.
EN_zzOFZSRqeMc_W000124 · in -17.6 dBFS · gain -2.4 dB · emolia-01998
(emotional numbness, teasing, embarrassment·measured, frequent disfluency, light breath, casual)(low mumble) Uh, so maybe it's already there. It's just built into the metric we just, (ahem) uh.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, teasing, embarrassment; style: casual, monologue; average recording, no background noise; genuineness 2.9/6; vocal-burst blend 1.0/10; 5.4s, EN.
EN_zzOFZSRqeMc_W000125 · in -18.3 dBFS · gain -1.7 dB · emolia-01998
Intoxication Altered States of Consciousness ↑ (unconstrained axis: Sexual Lust)identity −0.11emotion 90 % emotion__B1__T0.20__C0.25__INTERNAL · #13
This chain comes from the one-sided rule: only Intoxication Altered States of Consciousness had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Intoxication Altered States of Consciousness clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.25.
Nothing was asked of the other axis, and in fact Sexual Lust drifts down from 0.99 to 0.85 (-0.14), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.07, then +0.18 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 42 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.714 before conversion and 0.601 after — it fell by 0.113. Neighbour-to-neighbour the worst pair went 0.733 → 0.647. (The earlier render, with segment 1 left raw, scores 0.593 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.245 in the original and +0.221 after conversion — 90 % of the delta retained, which is essentially all of it. On the other named axis, Sexual Lust, -0.140 became -0.132.
Quality. Mean predicted overall quality across the segments went 2.79 → 3.15 (+0.36) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.714 → 0.601-0.113identity cos neighbours 0.733 → 0.647d_b rescored +0.245 → +0.221d_a rescored -0.140 → -0.132d_a mined -0.142d_b mined 0.245min_cos_consec (site) 0.8099min_cos_anchor (site) 0.8099dataset podcastlang enspeaker 484482total 41.6schain gain +2.9 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, fairly steady, light breath
(sexual lust, relief, contemplation · normal-paced, normally alert, slightly relaxed, casual)And when you shut down, by definition, that means you (ahem) cannot connect. You connect less and less and less.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as sexual lust, relief, contemplation; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 2.8/6; vocal-burst blend 3.1/10; 6.5s, EN.
484482_00099448 · in -25.4 dBFS · gain +5.4 dB · podcast-02663
(contemplation, fear· normal-paced, normally alert, neutral tension, casual)Now you still may be having, you know, physical intimacy of various forms, but it's not going to be anywhere near as fulfilling.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as contemplation, fear; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.2/6; vocal-burst blend 4.3/10; 7.8s, EN.
484482_00100100 · in -25.3 dBFS · gain +5.3 dB · podcast-02666
(intoxication altered states of consciousness, amusement, fatigue exhaustion·measured, very low-energy, relaxed, casual)a chore, an entitlement or (low mumble) uh yeah, depending on which gender you are. And (low mumble) uh and and it also becomes just basically a physical release. That's that's it. You know, I I started feeling frisky and you know I need a release and uh (low mumble) so let's go do it. And that (ahem) oh God that is sad. You know, and then you have the baby boomers as you said that (low mumble) uh for many of them uh
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as intoxication altered states of consciousness, amusement, fatigue exhaustion; style: casual, conversational; below-average recording, quiet background; genuineness 3.9/6; vocal-burst blend 5.0/10; 27.6s, EN.
484482_00101336 · in -25.1 dBFS · gain +5.1 dB · podcast-04428
This chain comes from the one-sided rule: only Longing had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Longing clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.21.
Nothing was asked of the other axis, and in fact Infatuation drifts down from 0.91 to 0.83 (-0.08), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.15, then +0.03, then -0.00, then +0.03 — not a clean run: step 3 moves back the other way by 0.00 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.78 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.78 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 40 s · ja · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.748 before conversion and 0.680 after — it fell by 0.068. Neighbour-to-neighbour the worst pair went 0.737 → 0.772. (The earlier render, with segment 1 left raw, scores 0.728 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
The emotional move did not survive. Re-scored end to end, Longing moved +0.207 in the original and -0.050 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Infatuation, -0.077 became -0.603.
Quality. Mean predicted overall quality across the segments went 3.11 → 3.14 (+0.02) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.748 → 0.680-0.068identity cos neighbours 0.737 → 0.772d_b rescored +0.207 → -0.050d_a rescored -0.077 → -0.603d_a mined -0.077d_b mined 0.207min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang jaspeaker JA_B00002_S08546total 38.4schain gain -0.2 dBseam step 1.6 dBcrossfades 150/100/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · balanced body, no background noise, measured, slightly relaxed, no disfluency, clear, light breath
This chain comes from the one-sided rule: only Hope Enthusiasm Optimism had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Hope Enthusiasm Optimism at the very top of the corpus — 0.99, higher than 99 % of clips in this corpus — and works its way down to clearly present at 0.73, higher than 73 % of clips in this corpus. That is a total fall of 0.26.
Nothing was asked of the other axis, and in fact Interest drifts down from 1.00 to 0.70 (-0.29), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are -0.10, then -0.16 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.73 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.71 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.73, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 26 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.747 before conversion and 0.803 after — it rose by 0.056. Neighbour-to-neighbour the worst pair went 0.637 → 0.762. (The earlier render, with segment 1 left raw, scores 0.722 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved -0.261 in the original and -0.358 after conversion — 137 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Interest, -0.292 became -0.419.
Quality. Mean predicted overall quality across the segments went 2.99 → 3.03 (+0.05) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.747 → 0.803+0.056identity cos neighbours 0.637 → 0.762d_b rescored -0.261 → -0.358d_a rescored -0.292 → -0.419d_a mined -0.293d_b mined -0.260min_cos_consec (site) 0.7115min_cos_anchor (site) 0.7328dataset emolialang enspeaker EN_B00056_S06084total 25.8schain gain +3.7 dBseam step 1.2 dBcrossfades 100/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normally alert, slightly relaxed, fairly steady
(interest, hope enthusiasm optimism, pride · brisk, casual, conversational)Then I can say, hey, how can I get stronger without getting bigger? And boom, I look towards powerlifting concepts. How can I get more powerful? How can I get faster?
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as interest, hope enthusiasm optimism, pride; style: casual, conversational; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 3.3/10; 8.1s, EN.
EN_B00056_S06084_W000223 · in -20.5 dBFS · gain +0.5 dB · emolia-01323
(concentration, interest ·normal-paced, casual, conversational)But I don't, uh, you (low mumble) know, again, want to lose fat. Okay, great. Or if I want physique changes. So we have all these different areas we can pick and choose from, uh, (low mumble) that have expertise in specific adaptations and develop ourselves perfect protocols, (ahem) uh, based on that information.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, interest; style: casual, conversational; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 6.0/10; 14.8s, EN.
EN_B00056_S06084_W000224 · in -21.1 dBFS · gain +1.1 dB · emolia-01323
(brisk, casual, conversational)Alright, the very first one we want to talk about is movement skill.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 3.4/6; vocal-burst blend 2.2/10; 3.3s, EN.
EN_B00056_S06084_W000225 · in -19.9 dBFS · gain -0.1 dB · emolia-01323
Intoxication Altered States of Consciousness ↑ (unconstrained axis: Awe)identity −0.21emotion 91 % emotion__B1__T0.20__C0.25__INTERNAL · #16
This chain comes from the one-sided rule: only Intoxication Altered States of Consciousness had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Intoxication Altered States of Consciousness below average — 0.31, lower than 69 % of clips in this corpus — and ends with it strongly present at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.46.
Nothing was asked of the other axis, and in fact Awe drifts down from 0.98 to 0.47 (-0.51), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.05, then +0.19, then +0.00, then +0.23 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.87 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.87 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 25 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.874 before conversion and 0.663 after — it fell by 0.211. Neighbour-to-neighbour the worst pair went 0.871 → 0.694. (The earlier render, with segment 1 left raw, scores 0.579 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.462 in the original and +0.420 after conversion — 91 % of the delta retained, which is essentially all of it. On the other named axis, Awe, -0.510 became -0.501.
Quality. Mean predicted overall quality across the segments went 2.75 → 2.82 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.874 → 0.663-0.211identity cos neighbours 0.871 → 0.694d_b rescored +0.462 → +0.420d_a rescored -0.510 → -0.501d_a mined -0.510d_b mined 0.462min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00036_S08918total 23.7schain gain +0.7 dBseam step 1.4 dBcrossfades 100/100/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(awe · measured, fairly steady, monologue, formal)Marina is a talented painter who creates beautiful landscapes and portraits.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as awe; style: monologue, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.1/10; 5.3s, ZH.
ZH_B00036_S08918_W000000 · in -19.2 dBFS · gain -0.8 dB · emolia-03640
(awe, affection· measured, fairly steady, monologue, formal)Marina is a talented painter who creates beautiful landscapes and portraits.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as awe, affection; style: monologue, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.0/10; 5.2s, ZH.
ZH_B00036_S08918_W000001 · in -19.1 dBFS · gain -0.8 dB · emolia-03640
(normal-paced, steady, formal, monologue)The geologist identified the composition of different types of rocks.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.2/10; 4.5s, ZH.
ZH_B00036_S08918_W000002 · in -19.4 dBFS · gain -0.6 dB · emolia-03640
(normal-paced, fairly steady, formal, monologue)The geologist identified the composition of different types of rocks.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.3/10; 4.5s, ZH.
ZH_B00036_S08918_W000003 · in -19.5 dBFS · gain -0.5 dB · emolia-03640
(measured, steady, monologue, formal)The film directors creative vision brought the story to life on the screen.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 3.0/10; 4.8s, ZH.
ZH_B00036_S08918_W000004 · in -18.8 dBFS · gain -1.2 dB · emolia-03640
This chain comes from the one-sided rule: only Impatience and Irritability had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Impatience and Irritability strongly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.24.
Nothing was asked of the other axis, and in fact Fear drifts down from 1.00 to 0.45 (-0.55), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.08, then -0.04, then +0.19 — not a clean run: step 2 moves back the other way by 0.04 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.53 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.55 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.53, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 35 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.489 before conversion and 0.740 after — it rose by 0.251. Neighbour-to-neighbour the worst pair went 0.586 → 0.740. (The earlier render, with segment 1 left raw, scores 0.614 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.234 in the original and +0.441 after conversion — 189 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Fear, -0.552 became -0.550.
Quality. Mean predicted overall quality across the segments went 2.61 → 3.01 (+0.40) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.489 → 0.740+0.251identity cos neighbours 0.586 → 0.740d_b rescored +0.234 → +0.441d_a rescored -0.552 → -0.550d_a mined -0.552d_b mined 0.238min_cos_consec (site) 0.5544min_cos_anchor (site) 0.5327dataset podcastlang enspeaker 824473total 34.4schain gain +2.5 dBseam step 0.3 dBcrossfades 100/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, normally alert, moderately variable, average clarity, wide pitch range
(fear, amusement, embarrassment · normal-paced, neutral tension, some disfluency, casual)Personally, that song scared me. I'll be honest. It was a bit scary to see and to watch and to hear.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as fear, amusement, embarrassment; style: casual, conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 5.4/10; 5.2s, EN.
824473_00166912 · in -15.0 dBFS · gain -5.0 dB · podcast-00999
(disgust, infatuation, sexual lust· normal-paced, relaxed, some disfluency, casual)It's one of those things that you do when you're so waved, and you're like, oh, listen,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as disgust, infatuation, sexual lust; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 3.5/6; vocal-burst blend 4.1/10; 4.8s, EN.
824473_00167448 · in -18.3 dBFS · gain -1.7 dB · podcast-00973
(doubt, confusion, teasing· normal-paced, neutral tension, frequent disfluency, casual)but I don't know. I don't see I I see JT being the the talent in the group. Yeah, (ahem) but apparently um JT. Oh, so you think if JT's lips are definitely being written, then no
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as doubt, confusion, teasing; style: casual, conversational; average recording, quiet background; genuineness 5.5/6; vocal-burst blend 5.9/10; 9.8s, EN.
824473_00169112 · in -20.6 dBFS · gain +0.6 dB · podcast-01016
(impatience and irritability, jealousy and envy, doubt ·brisk, neutral tension, some disfluency, casual)I don't I've come to terms with the fact that people are getting their lyrics written. They probably have been for for a long time, but it's only now because social media, it's become a known thing. But yeah, you know, shout out to your Miami. She's gonna make her bread off it, she's gonna buy herself a new watch, and she's gonna keep it stepping. No,
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as impatience and irritability, jealousy and envy, doubt; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 10.0/10; 15.1s, EN.
824473_00171408 · in -18.8 dBFS · gain -1.2 dB · podcast-00975
This chain comes from the one-sided rule: only Infatuation had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Infatuation around average — 0.51, right about the corpus median — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.39.
Nothing was asked of the other axis, and in fact Sourness drifts down from 0.83 to 0.32 (-0.51), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.17, then +0.11, then +0.11 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.92 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.92 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 37 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.883 before conversion and 0.854 after — it fell by 0.029. Neighbour-to-neighbour the worst pair went 0.803 → 0.739. (The earlier render, with segment 1 left raw, scores 0.661 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.391 in the original and +0.433 after conversion — 111 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Sourness, -0.508 became -0.795.
Quality. Mean predicted overall quality across the segments went 2.96 → 3.08 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.883 → 0.854-0.029identity cos neighbours 0.803 → 0.739d_b rescored +0.391 → +0.433d_a rescored -0.508 → -0.795d_a mined -0.508d_b mined 0.391min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_ucqJiFnVtaktotal 35.8schain gain +0.2 dBseam step 0.8 dBcrossfades 100/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, measured, normally alert
(moderate pitch range, formal, newsreading)The course content of all the above courses is of international standard updating every three years
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 8.2s, EN.
EN_ucqJiFnVtak_W000287 · in -14.9 dBFS · gain -5.1 dB · emolia-01276
(concentration, triumph·fairly narrow pitch, formal, newsreading)As per the scheme of the Government of India entitled New Millennium Indian Technology Leadership Initiative
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, triumph; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 14.2s, EN.
EN_ucqJiFnVtak_W000288 · in -14.3 dBFS · gain -5.7 dB · emolia-01276
(pride·moderate pitch range, formal, monologue)The Council for Scientific and Industrial Research
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.1/10; 5.4s, EN.
EN_ucqJiFnVtak_W000289 · in -15.2 dBFS · gain -4.8 dB · emolia-01276
(fairly narrow pitch, formal, authoritative)Delhi, Indian Institute of Tropical Meteorology, Pune, Indian Institute of Science, Bangalore
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 8.4s, EN.
EN_ucqJiFnVtak_W000291 · in -14.8 dBFS · gain -5.2 dB · emolia-01276
Fear ↑ (unconstrained axis: Contemplation)identity −0.06emotion REVERSED emotion__B1__T0.20__C0.25__INTERNAL · #19
This chain comes from the one-sided rule: only Fear had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Fear around average — 0.54, higher than 54 % of clips in this corpus — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.29.
Nothing was asked of the other axis, and in fact Contemplation drifts down from 0.73 to 0.29 (-0.44), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.06, then -0.02, then +0.00, then +0.25 — not a clean run: step 2 moves back the other way by 0.02 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 26 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.833 before conversion and 0.770 after — it fell by 0.063. Neighbour-to-neighbour the worst pair went 0.815 → 0.770. (The earlier render, with segment 1 left raw, scores 0.746 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
The emotional move did not survive. Re-scored end to end, Fear moved +0.293 in the original and -0.170 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Contemplation, -0.438 became -0.187.
Quality. Mean predicted overall quality across the segments went 2.88 → 3.09 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.833 → 0.770-0.063identity cos neighbours 0.815 → 0.770d_b rescored +0.293 → -0.170d_a rescored -0.438 → -0.187d_a mined -0.438d_b mined 0.293min_cos_consec (site) 0.8375min_cos_anchor (site) 0.8375dataset emolialang zhspeaker ZH_B00021_S09318total 24.5schain gain +0.6 dBseam step 1.1 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady
This chain comes from the one-sided rule: only Infatuation had to get where it was going, by at least 0.20. The other emotion was left completely free.
The chain starts with Infatuation clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 92 % of clips in this corpus. That is a total rise of 0.21.
Nothing was asked of the other axis, and in fact Confusion barely moves at all, sitting near 0.92 throughout.
It takes 2 clips to get there. Clip to clip the moves are +0.21 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.87 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.87 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 20 s · zh · emolia
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.857 before conversion and 0.850 after — it fell by 0.007. Neighbour-to-neighbour the worst pair went 0.857 → 0.850. (The earlier render, with segment 1 left raw, scores 0.777 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.209 in the original and +0.089 after conversion — 42 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Confusion, -0.004 became +0.067.
Quality. Mean predicted overall quality across the segments went 3.01 → 3.27 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.857 → 0.850-0.007identity cos neighbours 0.857 → 0.850d_b rescored +0.209 → +0.089d_a rescored -0.004 → +0.067d_a mined -0.005d_b mined 0.209min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00036_S09151total 19.2schain gain +1.4 dBseam step 0.0 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-bright, fairly smooth, balanced body, average recording, no background noise, normally alert, slightly relaxed, fairly steady