Manifest tier. emotion, rule B1, T=0.7, step cap 0.25. Population 16,409 chains (212 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 13,485.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words
are a real non-speech sound, printed where it happens.
The full generated caption for any clip is under “full caption & clip
details”. Its perceived-gender and background-noise clauses were re-rendered from
the numeric buckets, because the versions stored in the corpus index had those two
ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
The tags and captions describe the ORIGINAL clips — they are the
corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the
converted audio are the ones printed in each card's own paragraph and chips.
This chain comes from the one-sided rule: only Hope Enthusiasm Optimism had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Hope Enthusiasm Optimism barely there — 0.20, lower than 80 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.74.
Nothing was asked of the other axis, and in fact Emotional Numbness drifts down from 0.89 to 0.44 (-0.44), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.18, then +0.19, then +0.25, then +0.12 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.65 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.65 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 60 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.398 before conversion and -0.005 after — it fell by 0.403. Neighbour-to-neighbour the worst pair went 0.610 → 0.209. (The earlier render, with segment 1 left raw, scores 0.330 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.738 in the original and +0.702 after conversion — 95 % of the delta retained, which is essentially all of it. On the other named axis, Emotional Numbness, -0.443 became -0.403.
Quality. Mean predicted overall quality across the segments went 2.83 → 3.05 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.398 → -0.005-0.403identity cos neighbours 0.610 → 0.209d_b rescored +0.738 → +0.702d_a rescored -0.443 → -0.403d_a mined -0.443d_b mined 0.738min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00060_S09680total 58.3schain gain +5.2 dBseam step 2.4 dBcrossfades 100/100/150/100 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, quiet background, fairly steady
(measured, subdued, slightly relaxed, casual)So, (low mumble) uh, as we mentioned before,
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual; below-average recording, quiet background; genuineness 4.1/6; vocal-burst blend 1.1/10; 3.8s, EN.
EN_B00060_S09680_W000001 · in -16.9 dBFS · gain -3.0 dB · emolia-01389
(normal-paced, normally alert, slightly relaxed, monologue)Okay, so for detail, as I mentioned, go back to that Linux device driver book. Then, (low mumble) uhm, no more details.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 2.3/10; 7.5s, EN.
EN_B00060_S09680_W000002 · in -17.8 dBFS · gain -2.2 dB · emolia-01389
(normal-paced, normally alert, slightly relaxed, didactic)So another part is our process control. This is a huge part. Basically you can think about
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.4/10; 6.4s, EN.
EN_B00060_S09680_W000003 · in -18.1 dBFS · gain -1.9 dB · emolia-01389
(concentration·measured, normally alert, neutral tension, casual)Start process this IO request. Use our start request. You can see here. Based on this request. Then from this here, then we do real operation.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: casual, monologue; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 2.3/10; 11.3s, EN.
EN_B00060_S09680_W000004 · in -20.5 dBFS · gain +0.5 dB · emolia-01389
(hope enthusiasm optimism, relief· measured, subdued, slightly relaxed, monologue)(ahem) Generally speaking, we need (ahem) our program, which consists of instructions and (ahem) data, right? So we need a program and data in order to run our process. Then our program and the data will be put into memory. So somehow we need to find some physical memory space to hold our instruction and the data.
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism, relief; style: monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 4.4/10; 30.0s, EN.
EN_B00060_S09680_W000005 · in -20.0 dBFS · gain +0.0 dB · emolia-01389
This chain comes from the one-sided rule: only Infatuation had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Infatuation barely there — 0.23, lower than 77 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.73.
Nothing was asked of the other axis, and in fact Concentration barely moves at all, sitting near 0.85 throughout.
It takes 5 clips to get there. Clip to clip the moves are +0.13, then +0.21, then +0.20, then +0.19 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.80 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.80 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 31 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.490 before conversion and 0.565 after — it rose by 0.076. Neighbour-to-neighbour the worst pair went 0.564 → 0.529. (The earlier render, with segment 1 left raw, scores 0.433 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.730 in the original and +0.445 after conversion — 61 % of the delta retained. On the other named axis, Concentration, -0.042 became +0.018.
Quality. Mean predicted overall quality across the segments went 2.80 → 2.92 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.490 → 0.565+0.076identity cos neighbours 0.564 → 0.529d_b rescored +0.730 → +0.445d_a rescored -0.042 → +0.018d_a mined -0.043d_b mined 0.730min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00062_S08793total 29.5schain gain +3.0 dBseam step 2.4 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normally alert, slightly relaxed, average clarity
(normal-paced, fairly steady, some disfluency, monologue)So one thing I could do is just what I typed, (low mumble) uh, one, or what I wrote before. One divided by sine of 80 degrees.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 0.9/10; 7.5s, EN.
EN_B00062_S08793_W000004 · in -18.0 dBFS · gain -2.0 dB · emolia-01431
(emotional numbness·measured, steady, frequent disfluency, monologue)Sign of 30 degrees, sign is opposite over our potnus.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: monologue, formal; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 0.4/10; 4.5s, EN.
EN_B00062_S08793_W000005 · in -18.4 dBFS · gain -1.6 dB · emolia-01431
(teasing, amusement·normal-paced, fairly steady, some disfluency, casual)So one thing, if I look at this angle measure, I know that A plus B has to be 90. So, looks like angle B must be 60 degrees.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as teasing, amusement; style: casual, conversational; good recording, quiet background; genuineness 1.6/6; vocal-burst blend 2.1/10; 6.8s, EN.
EN_B00062_S08793_W000006 · in -18.7 dBFS · gain -1.3 dB · emolia-01431
(normal-paced, fairly steady, little disfluency, monologue)So what I'd like to do is get a couple other (ahem) trig functions defined up here.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 1.8/10; 3.7s, EN.
EN_B00062_S08793_W000007 · in -19.2 dBFS · gain -0.8 dB · emolia-01431
(infatuation, emotional numbness·measured, steady, frequent disfluency, monologue)And then same thing here, gonna rationalize that denominator, so this would be 2 root 5 over 5.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation, emotional numbness; style: monologue, didactic; good recording, quiet background; genuineness 1.6/6; vocal-burst blend 0.4/10; 7.9s, EN.
EN_B00062_S08793_W000008 · in -22.0 dBFS · gain +2.0 dB · emolia-01431
This chain comes from the one-sided rule: only Disgust had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Disgust barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.75.
Nothing was asked of the other axis, and in fact Pride drifts down from 0.96 to 0.53 (-0.44), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.19, then +0.20, then +0.12, then +0.24 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.89 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.89 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 51 s · ja · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.777 before conversion and 0.796 after — it rose by 0.019. Neighbour-to-neighbour the worst pair went 0.777 → 0.743. (The earlier render, with segment 1 left raw, scores 0.582 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disgust moved +0.752 in the original and +0.322 after conversion — 43 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Pride, -0.437 became -0.381.
Quality. Mean predicted overall quality across the segments went 2.89 → 3.19 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.777 → 0.796+0.019identity cos neighbours 0.777 → 0.743d_b rescored +0.752 → +0.322d_a rescored -0.437 → -0.381d_a mined -0.437d_b mined 0.752min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang jaspeaker JA_B00004_S05490total 50.1schain gain +2.1 dBseam step 0.3 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, balanced body, quiet background, normally alert, slightly relaxed, fairly steady, moderate pitch range
This chain comes from the one-sided rule: only Pain had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Pain barely there — 0.20, lower than 80 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.76.
Nothing was asked of the other axis, and in fact Relief drifts down from 0.87 to 0.64 (-0.23), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.15, then +0.20, then +0.17, then +0.24 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.85 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.85 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 42 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.693 before conversion and 0.684 after — it fell by 0.010. Neighbour-to-neighbour the worst pair went 0.795 → 0.723. (The earlier render, with segment 1 left raw, scores 0.577 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.759 in the original and +0.389 after conversion — 51 % of the delta retained. On the other named axis, Relief, -0.241 became -0.321.
Quality. Mean predicted overall quality across the segments went 2.55 → 2.85 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.693 → 0.684-0.010identity cos neighbours 0.795 → 0.723d_b rescored +0.759 → +0.389d_a rescored -0.241 → -0.321d_a mined -0.229d_b mined 0.759min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_QQsu0wCQbXytotal 40.8schain gain +3.7 dBseam step 1.9 dBcrossfades 150/150/100/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, balanced body, slightly relaxed, moderate pitch range, light breath
(normal-paced, normally alert, fairly steady, casual)Once your old beliefs are out of the way, you can begin to analyze your situation.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: casual, authoritative; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 1.0/10; 4.2s, EN.
EN_QQsu0wCQbXy_W000049 · in -18.9 dBFS · gain -1.1 dB · emolia-00912
(emotional numbness·measured, normally alert, fairly steady, casual)And what you do will flow logically from that.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as emotional numbness; style: casual, monologue; average recording, no background noise; genuineness 1.8/6; vocal-burst blend 1.6/10; 3.4s, EN.
EN_QQsu0wCQbXy_W000050 · in -19.6 dBFS · gain -0.4 dB · emolia-00912
(hope enthusiasm optimism·normal-paced, normally alert, steady, monologue)Some radicals envision a free world and work toward that.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as hope enthusiasm optimism; style: monologue, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.6/10; 4.2s, EN.
EN_QQsu0wCQbXy_W000051 · in -19.6 dBFS · gain -0.4 dB · emolia-00912
(normal-paced, energised, fairly steady, monologue)Others think the world of the future is so hard to predict there's no point in working toward a specific long-term goal, and we can make it up as we go along.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: monologue, dramatic; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.9/10; 9.0s, EN.
EN_QQsu0wCQbXy_W000052 · in -19.0 dBFS · gain -1.0 dB · emolia-00912
(contemplation, pain, interest· normal-paced, normally alert, fairly steady, monologue)There are two perspectives on achieving a world without oppressive institutions, where it's impossible for anyone to simply enter a system or acquire some resources and then go out and enslave people. That world is a long way away, but it's worth fighting for. There are many philosophies of liberation.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as contemplation, pain, interest; style: monologue, dramatic; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 1.2/10; 20.8s, EN.
EN_QQsu0wCQbXy_W000053 · in -19.9 dBFS · gain -0.1 dB · emolia-00912
This chain comes from the one-sided rule: only Thankfulness Gratitude had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Thankfulness Gratitude barely there — 0.19, lower than 81 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.77.
Nothing was asked of the other axis, and in fact Concentration drifts down from 0.94 to 0.80 (-0.15), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.23, then +0.22, then +0.12 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 48 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.781 before conversion and 0.797 after — it rose by 0.016. Neighbour-to-neighbour the worst pair went 0.636 → 0.609. (The earlier render, with segment 1 left raw, scores 0.698 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.795 in the original and +0.780 after conversion — 98 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.146 became -0.086.
Quality. Mean predicted overall quality across the segments went 2.84 → 3.07 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.781 → 0.797+0.016identity cos neighbours 0.636 → 0.609d_b rescored +0.795 → +0.780d_a rescored -0.146 → -0.086d_a mined -0.147d_b mined 0.774min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_zGa7mazWuCytotal 46.6schain gain +3.4 dBseam step 1.2 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · quiet background, measured, normally alert, slightly relaxed
(concentration · steady, frequent disfluency, somewhat unclear, didactic)This is for each i, so this relation if I concoct for all i, can be written in the matrix form as H v is equal to
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: didactic, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.0/10; 10.4s, EN.
EN_zGa7mazWuCy_W000211 · in -15.4 dBFS · gain -4.6 dB · emolia-02479
(concentration ·moderately variable, frequent disfluency, slurred, didactic)U lambda to the power half. What is lambda to the power half? Lambda to the power half is simply the diagonal matrix, which is lambda 1 to the power half, lambda 2 to the power half, and lambda n to the power half, the square root of the diagonal matrix lambda.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration; style: didactic, monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 2.5/10; 17.4s, EN.
EN_zGa7mazWuCy_W000212 · in -19.2 dBFS · gain -0.8 dB · emolia-02479
(fairly steady, frequent disfluency, average clarity, didactic)V is orthogonal, so I can multiply both sides by V transpose.
full caption & clip details
A child masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, thin; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: didactic, casual; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.7/10; 5.5s, EN.
EN_zGa7mazWuCy_W000213 · in -13.8 dBFS · gain -6.2 dB · emolia-02479
(jealousy and envy·moderately variable, some disfluency, average clarity, didactic)If I multiply both sides by v transpose, because v v transpose is i, I get this relation. Look at this now.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as jealousy and envy; style: didactic, authoritative; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 0.9/10; 7.6s, EN.
EN_zGa7mazWuCy_W000214 · in -16.2 dBFS · gain -3.8 dB · emolia-02479
(thankfulness gratitude·fairly steady, frequent disfluency, average clarity, monologue)I have now expressed the product of u lambda to the power half n v transpose and this
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as thankfulness gratitude; style: monologue, didactic; good recording, quiet background; genuineness 1.4/6; vocal-burst blend 0.1/10; 6.5s, EN.
EN_zGa7mazWuCy_W000215 · in -16.7 dBFS · gain -3.3 dB · emolia-02479
This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Concentration barely there — 0.08, lower than 92 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.83.
Nothing was asked of the other axis, and in fact Fatigue Exhaustion drifts down from 0.99 to 0.12 (-0.86), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.24, then +0.14, then +0.21 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.72 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.72 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 32 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.597 before conversion and 0.565 after — it fell by 0.032. Neighbour-to-neighbour the worst pair went 0.600 → 0.633. (The earlier render, with segment 1 left raw, scores 0.474 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.831 in the original and +0.756 after conversion — 91 % of the delta retained, which is essentially all of it. On the other named axis, Fatigue Exhaustion, -0.864 became +0.097.
Quality. Mean predicted overall quality across the segments went 2.78 → 2.98 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.597 → 0.565-0.032identity cos neighbours 0.600 → 0.633d_b rescored +0.831 → +0.756d_a rescored -0.864 → +0.097d_a mined -0.864d_b mined 0.831min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00068_S06406total 30.1schain gain +2.9 dBseam step 3.2 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · balanced body, normally alert, fairly steady
(fatigue exhaustion · normal-paced, slightly relaxed, some disfluency, casual)One of my policies kept failing. So just double check it.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as fatigue exhaustion; style: casual, conversational; good recording, no background noise; genuineness 3.6/6; vocal-burst blend 3.7/10; 3.1s, EN.
EN_B00068_S06406_W000012 · in -24.3 dBFS · gain +4.3 dB · emolia-01552
(contentment·measured, slightly relaxed, frequent disfluency, whispered)Okay. Does that make sense? The domain is happy. IT are frowny.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, slightly guarded; reads as contentment; style: whispered, casual; good recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.8/10; 7.2s, EN.
EN_B00068_S06406_W000013 · in -23.4 dBFS · gain +3.4 dB · emolia-01552
(normal-paced, slightly relaxed, little disfluency, casual)It will prompt you and say, do you want to link the GPO that you've selected? Ask for confirmation and you can see it listed there.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, slightly rough, balanced body; clear, little disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, conversational; good recording, quiet background; genuineness 1.6/6; vocal-burst blend 1.0/10; 7.0s, EN.
EN_B00068_S06406_W000014 · in -21.7 dBFS · gain +1.7 dB · emolia-01552
(fatigue exhaustion, doubt, relief·measured, relaxed, some disfluency, casual)I'm probably gonna have to restart this, reboot it. But I'll let it try.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, submissive, neutral openness; reads as fatigue exhaustion, doubt, relief; style: casual, conversational; below-average recording, quiet background; mildly explicit content; genuineness 4.1/6; vocal-burst blend 1.1/10; 4.5s, EN.
EN_B00068_S06406_W000015 · in -23.5 dBFS · gain +3.5 dB · emolia-01552
(concentration·normal-paced, slightly relaxed, some disfluency, casual)Okay, section one, here we go. We're going to create some Active Directory, (low mumble) uh, objects. So we'll start with the Server Manager on DC01.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as concentration; style: casual, monologue; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.7/10; 9.1s, EN.
EN_B00068_S06406_W000016 · in -23.4 dBFS · gain +3.4 dB · emolia-01552
This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Emotional Numbness barely there — 0.13, lower than 87 % of clips in this corpus — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.70.
Nothing was asked of the other axis, and in fact Doubt drifts down from 0.72 to 0.37 (-0.35), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.22, then +0.15, then +0.14 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 32 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.758 before conversion and 0.763 after — it rose by 0.004. Neighbour-to-neighbour the worst pair went 0.758 → 0.763. (The earlier render, with segment 1 left raw, scores 0.649 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.704 in the original and +0.270 after conversion — 38 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Doubt, -0.352 became -0.387.
Quality. Mean predicted overall quality across the segments went 2.98 → 3.10 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.758 → 0.763+0.004identity cos neighbours 0.758 → 0.763d_b rescored +0.704 → +0.270d_a rescored -0.352 → -0.387d_a mined -0.352d_b mined 0.704min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00057_S06868total 30.4schain gain +0.2 dBseam step 1.3 dBcrossfades 100/150/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady
(measured, no disfluency, average clarity, authoritative)六个月之后,西家花也进入公司,跟斯维木一起经营。
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, monologue; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 2.9/10; 4.9s, ZH.
ZH_B00057_S06868_W000179 · in -21.7 dBFS · gain +1.7 dB · emolia-03850
(normal-paced, no disfluency, average clarity, formal)并在两千年正式成名。Zippo的CEO.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; average recording, no background noise; genuineness 1.5/6; vocal-burst blend 3.4/10; 3.7s, ZH.
ZH_B00057_S06868_W000180 · in -21.0 dBFS · gain +1.0 dB · emolia-03850
(measured, no disfluency, clear, didactic)谢家花后来露西以个人身份和通过自己控制的创投青蛙公司向zipps追加投资超过一千万美元。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, no background noise; genuineness 0.9/6; vocal-burst blend 1.1/10; 9.2s, ZH.
ZH_B00057_S06868_W000181 · in -22.2 dBFS · gain +2.2 dB · emolia-03850
(normal-paced, no disfluency, clear, formal)并引入了红杉资本妖四千四百万美元的投资。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 2.2/10; 4.1s, ZH.
ZH_B00057_S06868_W000182 · in -20.6 dBFS · gain +0.7 dB · emolia-03850
(measured, almost no disfluency, clear, didactic)Zipos的成功出售其创意这个投资。Zipoo的VC都借此赚得的盆满钵满,这个案例很精彩。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.2/10; 9.4s, ZH.
ZH_B00057_S06868_W000183 · in -21.7 dBFS · gain +1.7 dB · emolia-03850
This chain comes from the one-sided rule: only Hope Enthusiasm Optimism had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Hope Enthusiasm Optimism barely there — 0.20, lower than 80 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.70.
Nothing was asked of the other axis, and in fact Astonishment Surprise drifts down from 0.76 to 0.66 (-0.10), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.06, then +0.22, then +0.19 — a plateau around step 2, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 35 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.693 before conversion and 0.722 after — it rose by 0.029. Neighbour-to-neighbour the worst pair went 0.650 → 0.710. (The earlier render, with segment 1 left raw, scores 0.603 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.702 in the original and +0.695 after conversion — 99 % of the delta retained, which is essentially all of it. On the other named axis, Astonishment Surprise, -0.103 became -0.407.
Quality. Mean predicted overall quality across the segments went 2.98 → 3.17 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.693 → 0.722+0.029identity cos neighbours 0.650 → 0.710d_b rescored +0.702 → +0.695d_a rescored -0.103 → -0.407d_a mined -0.103d_b mined 0.702min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00002_S03399total 33.4schain gain +2.3 dBseam step 1.3 dBcrossfades 150/150/100/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, fairly steady
(fast, no disfluency, dramatic, authoritative)英国和美国等消费趋势型的经济体将继续面对巨大的贸易赤字。
full caption & clip details
A young adult feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: dramatic, authoritative; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 2.4/10; 5.8s, ZH.
ZH_B00002_S03399_W000015 · in -14.8 dBFS · gain -5.2 dB · emolia-03301
(normal-paced, no disfluency, authoritative, narration)赤子国家承受的巨大压力是毋庸置疑的。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, narration; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 6.2/10; 3.2s, ZH.
ZH_B00002_S03399_W000016 · in -14.5 dBFS · gain -5.5 dB · emolia-03301
(measured, almost no disfluency, authoritative, didactic)因为他们都是在承担着天量债务的国家,因而也必须实施自我调整的国家,而盈余国家显然没有这样的压力。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, didactic; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 2.2/10; 9.8s, ZH.
ZH_B00002_S03399_W000017 · in -15.9 dBFS · gain -4.0 dB · emolia-03301
(normal-paced, some disfluency, didactic, monologue)而且在这些国家中,很多都没有表现出通过扩大国内支出,让赤子国实现再平衡的欲望。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 3.1/10; 7.5s, ZH.
ZH_B00002_S03399_W000018 · in -15.1 dBFS · gain -4.9 dB · emolia-03301
(hope enthusiasm optimism, doubt·measured, no disfluency, monologue, didactic)至于各国为全球经济在平衡创造必要条件开展合作仪式,迄今为止尚未达成一致。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as hope enthusiasm optimism, doubt; style: monologue, didactic; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 2.4/10; 7.9s, ZH.
ZH_B00002_S03399_W000019 · in -14.8 dBFS · gain -5.2 dB · emolia-03301
This chain comes from the one-sided rule: only Doubt had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Doubt barely there — 0.22, lower than 78 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.73.
Nothing was asked of the other axis, and in fact Thankfulness Gratitude drifts down from 0.97 to 0.07 (-0.91), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.15, then +0.19, then +0.18, then +0.21 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 45 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.863 before conversion and 0.847 after — it fell by 0.016. Neighbour-to-neighbour the worst pair went 0.844 → 0.775. (The earlier render, with segment 1 left raw, scores 0.833 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.727 in the original and +0.701 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Thankfulness Gratitude, -0.905 became -0.883.
Quality. Mean predicted overall quality across the segments went 3.04 → 3.18 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.863 → 0.847-0.016identity cos neighbours 0.844 → 0.775d_b rescored +0.727 → +0.701d_a rescored -0.905 → -0.883d_a mined -0.905d_b mined 0.727min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00007_S00075total 44.1schain gain +2.0 dBseam step 2.1 dBcrossfades 100/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-bright, fairly smooth, no background noise, normally alert, slightly relaxed, fairly steady, clear, moderate pitch range
(thankfulness gratitude, relief, longing · normal-paced, some disfluency, monologue, dramatic)我们去看看发展中国家,城市化的典型者如墨西哥。据他们的学者说,百分之八十的城市化率至少比我们高百分之四十吧。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, relief, longing; style: monologue, dramatic; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 4.2/10; 9.2s, ZH.
ZH_B00007_S00075_W000304 · in -15.9 dBFS · gain -4.1 dB · emolia-03347
(thankfulness gratitude ·fast, no disfluency, monologue, formal)就是靠大型贫民窟去实现。而在大型贫民窟里面的就是这些大量的无地农民,他们土地被征占了,或者买卖了私有制条件下。
full caption & clip details
A young adult feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude; style: monologue, formal; average recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.5/10; 10.8s, ZH.
ZH_B00007_S00075_W000305 · in -15.8 dBFS · gain -4.2 dB · emolia-03347
(normal-paced, no disfluency, monologue, formal)任何大农场主都大规模占地,任何城里有钱人也都可以在乡下包一块地,至少弄个别墅。我们中国已经有一个亿的中产阶级人口,谁家手里没有几几十上百万。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 0.3/6; vocal-burst blend 2.8/10; 13.4s, ZH.
ZH_B00007_S00075_W000306 · in -15.1 dBFS · gain -4.9 dB · emolia-03347
(fast, some disfluency, casual, storytelling)但幸亏不仅仅是我说的。大家知道,在中央电视台的颁奖晚会上,王小丫说。
full caption & clip details
A young adult feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, storytelling; average recording, no background noise; genuineness 2.1/6; vocal-burst blend 4.0/10; 6.0s, ZH.
ZH_B00007_S00075_W000307 · in -15.3 dBFS · gain -4.7 dB · emolia-03347
(doubt· fast, some disfluency, didactic, monologue)温铁军有一句话说,农民头上三把刀、上学、告状和医疗。
full caption & clip details
A young adult feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as doubt; style: didactic, monologue; average recording, no background noise; genuineness 1.8/6; vocal-burst blend 3.4/10; 5.4s, ZH.
ZH_B00007_S00075_W000308 · in -15.5 dBFS · gain -4.5 dB · emolia-03347
This chain comes from the one-sided rule: only Thankfulness Gratitude had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Thankfulness Gratitude barely there — 0.15, lower than 85 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.76.
Nothing was asked of the other axis, and in fact Emotional Numbness climbs from 0.84 to 0.91 (+0.07), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.12, then +0.25, then +0.24, then +0.16 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 31 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.678 before conversion and 0.748 after — it rose by 0.070. Neighbour-to-neighbour the worst pair went 0.678 → 0.737. (The earlier render, with segment 1 left raw, scores 0.646 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
The emotional move did not survive. Re-scored end to end, Thankfulness Gratitude moved +0.764 in the original and -0.059 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Emotional Numbness, +0.073 became +0.213.
Quality. Mean predicted overall quality across the segments went 2.84 → 3.04 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.678 → 0.748+0.070identity cos neighbours 0.678 → 0.737d_b rescored +0.764 → -0.059d_a rescored +0.073 → +0.213d_a mined 0.073d_b mined 0.764min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00008_S08128total 29.6schain gain +1.2 dBseam step 0.9 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, light breath
(measured, steady, no disfluency, didactic)已经向通货膨胀宣战,把利率推到了前所未有的高度。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, formal; average recording, no background noise; genuineness 1.7/6; vocal-burst blend 1.5/10; 5.6s, ZH.
ZH_B00008_S08128_W000033 · in -20.2 dBFS · gain +0.2 dB · emolia-03360
(normal-paced, fairly steady, no disfluency, formal)现金流从白银等投资中抽离出来。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, didactic; good recording, no background noise; genuineness 2.4/6; vocal-burst blend 1.8/10; 3.2s, ZH.
ZH_B00008_S08128_W000034 · in -20.6 dBFS · gain +0.6 dB · emolia-03360
(measured, steady, some disfluency, monologue)亨特家族持有的大量的白银需要花费很多钱,比如仓储的费用。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, whispered; average recording, no background noise; genuineness 2.0/6; vocal-burst blend 1.1/10; 5.8s, ZH.
ZH_B00008_S08128_W000035 · in -18.7 dBFS · gain -1.3 dB · emolia-03360
(measured, steady, some disfluency, monologue)一九八零年三月二十五日,形势发生了翻天地翻天覆地的一个变化。亨特家族的白银交易员八十。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, whispered; average recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.9/10; 9.2s, ZH.
ZH_B00008_S08128_W000036 · in -19.9 dBFS · gain -0.1 dB · emolia-03360
(thankfulness gratitude, emotional numbness· measured, steady, no disfluency, formal)后知国际金属公司合伙人需要一点三五亿美元的追加保证金时。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, emotional numbness; style: formal, monologue; average recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.4/10; 6.6s, ZH.
ZH_B00008_S08128_W000037 · in -20.3 dBFS · gain +0.3 dB · emolia-03360
This chain comes from the one-sided rule: only Pain had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Pain barely there — 0.11, lower than 89 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.80.
Nothing was asked of the other axis, and in fact Relief drifts down from 0.90 to 0.20 (-0.70), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.20, then +0.17, then +0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.86 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.86 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 31 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.816 before conversion and 0.809 after — it fell by 0.008. Neighbour-to-neighbour the worst pair went 0.777 → 0.748. (The earlier render, with segment 1 left raw, scores 0.677 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.802 in the original and +0.436 after conversion — 54 % of the delta retained. On the other named axis, Relief, -0.697 became -0.443.
Quality. Mean predicted overall quality across the segments went 2.83 → 3.06 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.816 → 0.809-0.008identity cos neighbours 0.777 → 0.748d_b rescored +0.802 → +0.436d_a rescored -0.697 → -0.443d_a mined -0.697d_b mined 0.802min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00004_S03467total 30.1schain gain +2.4 dBseam step 1.1 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, average recording, normally alert, slightly relaxed, light breath
(measured, fairly steady, some disfluency, authoritative)所以呢心里面有点后悔,没有早点上来,但后悔肯定是没有用的,是吧?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, casual; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 2.1/10; 6.3s, ZH.
ZH_B00004_S03467_W000010 · in -18.3 dBFS · gain -1.7 dB · emolia-03319
(measured, steady, no disfluency, monologue)确认没有了以后呢,就订了酒店,把行李放下以后呢,出去享受一下大连的夜景。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 1.5/6; vocal-burst blend 1.2/10; 8.0s, ZH.
ZH_B00004_S03467_W000011 · in -19.3 dBFS · gain -0.7 dB · emolia-03319
(normal-paced, fairly steady, some disfluency, casual)另外呢,因为我当天的播音啊没有录。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 5.6/6; vocal-burst blend 3.4/10; 3.1s, ZH.
ZH_B00004_S03467_W000012 · in -17.1 dBFS · gain -2.9 dB · emolia-03319
(measured, fairly steady, some disfluency, monologue)但很多人呢也会让这些不顺利呢,让自己心情变得很糟糕。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 2.3/10; 5.4s, ZH.
ZH_B00004_S03467_W000013 · in -17.3 dBFS · gain -2.7 dB · emolia-03319
(pain· measured, fairly steady, some disfluency, monologue)而卡罗尔遇到的事情,要想保持积极的态度呢是非常困难的。苏勒博是夫妻呢,是在国外。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: monologue, didactic; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 3.7/10; 8.0s, ZH.
ZH_B00004_S03467_W000014 · in -18.8 dBFS · gain -1.2 dB · emolia-03319
This chain comes from the one-sided rule: only Pain had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Pain barely there — 0.11, lower than 89 % of clips in this corpus — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.73.
Nothing was asked of the other axis, and in fact Interest drifts down from 0.97 to 0.29 (-0.68), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.20, then +0.17, then +0.12 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 61 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.775 before conversion and 0.802 after — it rose by 0.027. Neighbour-to-neighbour the worst pair went 0.711 → 0.738. (The earlier render, with segment 1 left raw, scores 0.726 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.727 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Interest, -0.679 became -0.522.
Quality. Mean predicted overall quality across the segments went 2.93 → 3.12 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.775 → 0.802+0.027identity cos neighbours 0.711 → 0.738d_b rescored +0.727 → +0.000d_a rescored -0.679 → -0.522d_a mined -0.679d_b mined 0.727min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_ArAG7AtUfwgtotal 59.5schain gain +2.6 dBseam step 1.9 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, fairly steady, somewhat unclear, moderate pitch range
(interest, triumph · brisk, slightly relaxed, some disfluency, casual)Okay, yes, (surprised gasp) Maria has been kind of (ahem) instrumental in terms of working together with the HISP groups, (low mumble) uh, coordinating the project, kind of, (ahem) uh, you know, pushing together, (ahem) uh, some of the requirements which UNICEF have, but also some of the requirements which the country has.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, triumph; style: casual; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 5.9/10; 14.2s, EN.
EN_ArAG7AtUfwg_W000120 · in -19.1 dBFS · gain -0.9 dB · emolia-00685
(hope enthusiasm optimism, interest, elation·normal-paced, slightly relaxed, frequent disfluency, casual)(low mumble) Uhm, we have been working with the UNICEF team in terms of making sure that we also follow up some of these, (low mumble) uhm, you know, plans which we have (low mumble) along the way. So UNICEF has been cut off (ahem) quite instrumental.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as hope enthusiasm optimism, interest, elation; style: casual, monologue; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.3/10; 12.8s, EN.
EN_ArAG7AtUfwg_W000121 · in -20.1 dBFS · gain +0.1 dB · emolia-00685
(thankfulness gratitude· normal-paced, slightly relaxed, some disfluency, monologue)In terms of, (low mumble) uh, supporting us, but also kind of be part of the process as well. I think that has been kind of a key thing, (low mumble) uh, during this particular (low mumble) collaboration.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude; style: monologue, formal; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 1.4/10; 9.7s, EN.
EN_ArAG7AtUfwg_W000122 · in -21.1 dBFS · gain +1.1 dB · emolia-00685
(concentration·brisk, neutral tension, some disfluency, monologue)So this is kind of an example of a joint roadmap, (low mumble) uh, and a release plan, which we kind of had so much information right there. But, (ahem) uh, one thing which we kind of noticed that (ahem) we have kind of different layers of, (ahem) uh, (ahem) activities which we have. We have the implementation activities, uh, (low mumble) which the HIPs groups are more or less kind of supporting countries.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration; style: monologue; below-average recording, some background noise; genuineness 3.2/6; vocal-burst blend 5.9/10; 19.0s, EN.
EN_ArAG7AtUfwg_W000123 · in -18.4 dBFS · gain -1.6 dB · emolia-00685
(normal-paced, slightly relaxed, some disfluency, monologue)(low mumble) Uhm, into, you know, implementing the scorecard, BNA and action tracker.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.5/10; 4.5s, EN.
EN_ArAG7AtUfwg_W000124 · in -16.1 dBFS · gain -3.9 dB · emolia-00685
This chain comes from the one-sided rule: only Triumph had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Triumph essentially absent — 0.04, lower than 96 % of clips in this corpus — and ends with it strongly present at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.76.
Nothing was asked of the other axis, and in fact Contemplation drifts down from 0.96 to 0.71 (-0.26), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.13, then +0.22, then +0.20 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 54 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.785 before conversion and 0.805 after — it rose by 0.020. Neighbour-to-neighbour the worst pair went 0.758 → 0.783. (The earlier render, with segment 1 left raw, scores 0.677 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Triumph moved +0.761 in the original and +0.685 after conversion — 90 % of the delta retained, which is most of it. On the other named axis, Contemplation, -0.256 became -0.341.
Quality. Mean predicted overall quality across the segments went 2.97 → 3.16 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.785 → 0.805+0.020identity cos neighbours 0.758 → 0.783d_b rescored +0.761 → +0.685d_a rescored -0.256 → -0.341d_a mined -0.257d_b mined 0.761min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00000_S04501total 52.1schain gain +4.2 dBseam step 1.9 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · neutral-toned, fairly smooth, average recording, measured, slightly relaxed, fairly steady
This chain comes from the one-sided rule: only Sourness had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Sourness barely there — 0.17, lower than 83 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.81.
Nothing was asked of the other axis, and in fact Astonishment Surprise drifts down from 0.99 to 0.72 (-0.27), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.17, then +0.21, then +0.20 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of -0.15 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of -0.15 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 41 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.065 before conversion and 0.288 after — it rose by 0.223. Neighbour-to-neighbour the worst pair went 0.244 → 0.435. (The earlier render, with segment 1 left raw, scores 0.251 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sourness moved +0.813 in the original and +0.748 after conversion — 92 % of the delta retained, which is essentially all of it. On the other named axis, Astonishment Surprise, -0.331 became -0.237.
Quality. Mean predicted overall quality across the segments went 2.72 → 2.95 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.065 → 0.288+0.223identity cos neighbours 0.244 → 0.435d_b rescored +0.813 → +0.748d_a rescored -0.331 → -0.237d_a mined -0.267d_b mined 0.812min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_MLvd4-Nw3lototal 39.9schain gain +2.9 dBseam step 2.0 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · fairly smooth, quiet background, light breath
(astonishment surprise, affection · normal-paced, energised, slightly relaxed, monologue)We're all in this together. Suddenly became snitch on your neighbours.
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as astonishment surprise, affection; style: monologue, formal; good recording, quiet background; genuineness 0.6/6; vocal-burst blend 1.0/10; 5.0s, EN.
EN_MLvd4-Nw3lo_W000088 · in -16.5 dBFS · gain -3.5 dB · emolia-01719
(hope enthusiasm optimism, elation· normal-paced, normally alert, neutral tension, casual)And tell if people are, you know, breaking the rules and, you know, encouraging them to be, to, encouraging you to report them.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism, elation; style: casual, conversational; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 4.9/10; 8.3s, EN.
EN_MLvd4-Nw3lo_W000089 · in -18.8 dBFS · gain -1.2 dB · emolia-01719
(normal-paced, normally alert, slightly relaxed, monologue)And now in the concept of individual liberty has become comply.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: monologue, conversational; good recording, quiet background; genuineness 1.9/6; vocal-burst blend 1.4/10; 5.1s, EN.
EN_MLvd4-Nw3lo_W000090 · in -15.5 dBFS · gain -4.5 dB · emolia-01719
(disappointment, impatience and irritability, distress· normal-paced, normally alert, slightly relaxed, formal)Or down the line we will have to pay fines or be excluded from society.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as disappointment, impatience and irritability, distress; style: formal, monologue; good recording, quiet background; genuineness 1.8/6; vocal-burst blend 0.7/10; 5.2s, EN.
EN_MLvd4-Nw3lo_W000091 · in -18.3 dBFS · gain -1.7 dB · emolia-01719
(sourness, bitterness, malevolence malice·brisk, normally alert, neutral tension, casual)Just a quick one, you can see now with, uh, (low mumble) Harris leading out that people, (ahem) uh, have to self, uh, (low mumble) self isolate for two weeks if they come into the country and there's a thousand euro fine for non-compliance. So there's your first fine being rolled out, but it's against the people coming in. It's not against us, just against other people, but there's a fine there now.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as sourness, bitterness, malevolence malice; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 7.2/10; 17.1s, EN.
EN_MLvd4-Nw3lo_W000093 · in -18.1 dBFS · gain -1.9 dB · emolia-01719
This chain comes from the one-sided rule: only Impatience and Irritability had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Impatience and Irritability barely there — 0.18, lower than 82 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.76.
Nothing was asked of the other axis, and in fact Interest drifts down from 0.94 to 0.37 (-0.57), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.06, then +0.23, then +0.23, then +0.24 — a plateau around step 1, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.66 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.78 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.66, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 46 s · en · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.610 before conversion and 0.678 after — it rose by 0.068. Neighbour-to-neighbour the worst pair went 0.619 → 0.678. (The earlier render, with segment 1 left raw, scores 0.627 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.752 in the original and +0.717 after conversion — 95 % of the delta retained, which is essentially all of it. On the other named axis, Interest, -0.565 became -0.716.
Quality. Mean predicted overall quality across the segments went 2.72 → 2.89 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.610 → 0.678+0.068identity cos neighbours 0.619 → 0.678d_b rescored +0.752 → +0.717d_a rescored -0.565 → -0.716d_a mined -0.569d_b mined 0.765min_cos_consec (site) 0.7781min_cos_anchor (site) 0.6649dataset podcastlang enspeaker 24317total 45.1schain gain +3.6 dBseam step 1.3 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult feminine voice · neutral-toned, fairly smooth, balanced body, normally alert, some disfluency
(interest, sexual lust · normal-paced, slightly relaxed, moderately variable, conversational)looks (ahem) um so I'm gonna just describe it, Sister Marie Ann, and maybe let others share something they see, but it looks (ahem) um it's got like a tree type shape, but lots of connections. I just want you to look at it
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as interest, sexual lust; style: conversational, casual; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 4.8/10; 14.8s, EN.
24317_00311184 · in -19.7 dBFS · gain -0.3 dB · podcast-01760
(normal-paced, slightly relaxed, fairly steady, casual)I see the person too, Lisette. That was in the center.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 3.9/6; vocal-burst blend 3.2/10; 3.1s, EN.
24317_00313736 · in -21.6 dBFS · gain +1.6 dB · podcast-01784
(normal-paced, slightly relaxed, moderately variable, casual)When Sister Marie Ann and I had our pre-interview just to chat through, like what are the directions
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 0.7/10; 7.0s, EN.
24317_00369956 · in -19.6 dBFS · gain -0.4 dB · podcast-01838
(sourness, affection, doubt· normal-paced, neutral tension, moderately variable, casual)We are going to need to unpack this further. I knew so, and isn't that the beauty of a life of a journey that the unpacking is never done, right? That's the thing. That's the
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as sourness, affection, doubt; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 4.5/10; 14.1s, EN.
24317_00371232 · in -20.5 dBFS · gain +0.5 dB · podcast-01773
(impatience and irritability, astonishment surprise, disappointment·brisk, neutral tension, moderately variable, conversational)gift, finding light along the way, and that it's never we never get to where we think we're going.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as impatience and irritability, astonishment surprise, disappointment; style: conversational, casual; good recording, no background noise; genuineness 2.8/6; vocal-burst blend 3.6/10; 6.7s, EN.
24317_00372768 · in -19.7 dBFS · gain -0.3 dB · podcast-01755
This chain comes from the one-sided rule: only Embarrassment had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Embarrassment barely there — 0.23, lower than 77 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.71.
Nothing was asked of the other axis, and in fact Jealousy and Envy drifts down from 0.97 to 0.02 (-0.94), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.23, then +0.06, then +0.25 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 60 s · fr · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.703 before conversion and 0.608 after — it fell by 0.095. Neighbour-to-neighbour the worst pair went 0.721 → 0.641. (The earlier render, with segment 1 left raw, scores 0.503 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.722 in the original and +0.257 after conversion — 36 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Jealousy and Envy, -0.943 became -0.857.
Quality. Mean predicted overall quality across the segments went 2.96 → 3.06 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.703 → 0.608-0.095identity cos neighbours 0.721 → 0.641d_b rescored +0.722 → +0.257d_a rescored -0.943 → -0.857d_a mined -0.944d_b mined 0.711min_cos_consec (site) 0.8805min_cos_anchor (site) 0.8602dataset podcastlang frspeaker 164249total 58.7schain gain +0.9 dBseam step 2.2 dBcrossfades 150/150/100/100 ms
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · average recording, fairly steady
(jealousy and envy, anger, contemplation · measured, very low-energy, relaxed, casual)What's that jalous? (low mumble)
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as jealousy and envy, anger, contemplation; style: casual, monologue; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 9.6/10; 16.0s, FR.
164249_00320808 · in -30.3 dBFS · gain +10.3 dB · podcast-02438
(contemplation ·slow, very low-energy, relaxed, monologue)Or (low mumble) always we (low mumble) attend (low mumble) (low mumble) more of your relationship, for example. There are people who
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, narrow pitch range, normal breath; affect is mildly negative, submissive, slightly guarded; reads as contemplation; style: monologue, whispered; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 2.2/10; 16.0s, FR.
164249_00322412 · in -28.1 dBFS · gain +8.1 dB · podcast-02430
(emotional numbness, helplessness, shame·measured, normally alert, fully relaxed, casual)(low mumble) attend too of their relationship.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, fully relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, thin; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as emotional numbness, helplessness, shame; style: casual, whispered; average recording, no background noise; explicit content; genuineness 3.2/6; vocal-burst blend 6.1/10; 3.4s, FR.
164249_00324012 · in -28.7 dBFS · gain +8.7 dB · podcast-03205
(disappointment, sadness, contentment· measured, very low-energy, relaxed, monologue)(low mumble) (surprised gasp) Recently, Sandone who has said that the problem was the confiance. I think it's Sandrine that did.
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is slightly warm, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as disappointment, sadness, contentment; style: monologue, whispered; average recording, quiet background; mildly explicit content; genuineness 2.4/6; vocal-burst blend 8.3/10; 16.9s, FR.
164249_00325272 · in -29.2 dBFS · gain +9.2 dB · podcast-02461
(embarrassment, relief, sourness· measured, very low-energy, relaxed, casual)(ahem) Okay, identify your (low mumble)
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, submissive, neutral openness; reads as embarrassment, relief, sourness; style: casual, conversational; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 4.5/10; 6.9s, FR.
164249_00326956 · in -29.4 dBFS · gain +9.4 dB · podcast-03211
This chain comes from the one-sided rule: only Triumph had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Triumph essentially absent — 0.04, lower than 96 % of clips in this corpus — and ends with it strongly present at 0.78, higher than 78 % of clips in this corpus. That is a total rise of 0.74.
Nothing was asked of the other axis, and in fact Concentration drifts down from 0.94 to 0.42 (-0.52), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.20, then +0.19, then +0.14 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.87 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.87 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 73 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.733 before conversion and 0.641 after — it fell by 0.092. Neighbour-to-neighbour the worst pair went 0.826 → 0.727. (The earlier render, with segment 1 left raw, scores 0.543 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Triumph moved +0.743 in the original and +0.655 after conversion — 88 % of the delta retained, which is most of it. On the other named axis, Concentration, -0.522 became -0.563.
Quality. Mean predicted overall quality across the segments went 2.96 → 3.10 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.733 → 0.641-0.092identity cos neighbours 0.826 → 0.727d_b rescored +0.743 → +0.655d_a rescored -0.522 → -0.563d_a mined -0.522d_b mined 0.743min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_EloFEJpvF9Atotal 71.9schain gain +0.8 dBseam step 1.1 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(concentration · fairly steady, almost no disfluency, light breath, newsreading)Moore's law, which was formulated around 1965, calculated that the number of transistors in a dense integrated circuit doubles approximately every two years.The proliferation of the smaller and less expensive personal computers and improvements in computing power by the early 1980s resulted in sudden access to and the ability to share and store information for increasing numbers of workers.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 24.4s, EN.
EN_EloFEJpvF9A_W000010 · in -15.1 dBFS · gain -4.9 dB · emolia-00529
(fairly steady, no disfluency, light breath, formal)connectivity between computers within companies led to the ability of workers at different levels to access greater amounts of information
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.3/10; 8.1s, EN.
EN_EloFEJpvF9A_W000011 · in -14.7 dBFS · gain -5.3 dB · emolia-00529
(emotional numbness, pain· fairly steady, no disfluency, minimal breath, newsreading)The world's technological capacity to store information grew from 2.6' optimally compressed exabytes in 1986 to 15.8 in 1993, over 54.5 in 2000, and to 295' optimally compressed exabytes in 2007.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, pain; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 18.6s, EN.
EN_EloFEJpvF9A_W000013 · in -15.5 dBFS · gain -4.5 dB · emolia-00529
(emotional numbness ·steady, no disfluency, light breath, newsreading)This is the informational equivalent to less than 1730 Mbcd ROM per person in 1986 539 MB per person, roughly 4 CD ROM per person of 1993, 12 CD ROM per person in the year 2000.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 17.2s, EN.
EN_EloFEJpvF9A_W000014 · in -14.6 dBFS · gain -5.4 dB · emolia-00529
(fairly steady, no disfluency, light breath, formal)And almost 61 CD-ROM per person in 2007
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.7/10; 4.3s, EN.
EN_EloFEJpvF9A_W000015 · in -14.9 dBFS · gain -5.1 dB · emolia-00529
This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Concentration barely there — 0.16, lower than 84 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.74.
Nothing was asked of the other axis, and in fact Fatigue Exhaustion drifts down from 0.88 to 0.83 (-0.05), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.10, then +0.21, then +0.24, then +0.20 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 32 s · fr · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.755 before conversion and 0.792 after — it rose by 0.037. Neighbour-to-neighbour the worst pair went 0.838 → 0.769. (The earlier render, with segment 1 left raw, scores 0.651 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.737 in the original and +0.616 after conversion — 84 % of the delta retained, which is most of it. On the other named axis, Fatigue Exhaustion, -0.053 became -0.446.
Quality. Mean predicted overall quality across the segments went 2.92 → 3.01 (+0.09) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.755 → 0.792+0.037identity cos neighbours 0.838 → 0.769d_b rescored +0.737 → +0.616d_a rescored -0.053 → -0.446d_a mined -0.053d_b mined 0.737min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang frspeaker FR_BCoQVWd-9tktotal 30.7schain gain +0.2 dBseam step 0.5 dBcrossfades 100/100/100/100 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, normally alert, slightly relaxed, somewhat unclear
(measured, steady, some disfluency, formal)Et tu as cette autre vidéo de cours. Donc, ça fait neuf vidéos en tout. Ça fait à peu près
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; average recording, no background noise; genuineness 2.1/6; vocal-burst blend 0.1/10; 5.0s, FR.
FR_BCoQVWd-9tk_W000021 · in -15.9 dBFS · gain -4.1 dB · emolia-02638
(normal-paced, fairly steady, some disfluency, casual)Deux bonnes heures, c'est-à-dire à peu près deux heures trente de vidéos. Donc, il y a de quoi faire.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, formal; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 1.4/10; 3.5s, FR.
FR_BCoQVWd-9tk_W000022 · in -16.1 dBFS · gain -3.9 dB · emolia-02638
(measured, fairly steady, frequent disfluency)ça te fait un cours vraiment complet, donc tu me vois travailler le morceau, euh, (low mumble) (low mumble) euh, en même temps que toi.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.3/10; 5.5s, FR.
FR_BCoQVWd-9tk_W000023 · in -17.0 dBFS · gain -3.0 dB · emolia-02638
(measured, fairly steady, some disfluency, monologue)(surprised gasp) euh, tu n'as qu'à suivre mes indications et m'imiter pour apprendre un joli morceau. Et donc, si tu regardes euh,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 1.2/10; 6.2s, FR.
FR_BCoQVWd-9tk_W000024 · in -17.6 dBFS · gain -2.4 dB · emolia-02638
(measured, fairly steady, frequent disfluency, monologue)On va dire, allez, le premier jour, tu regardes la vidéo de présentation plus (low mumble) la vidéo sur la gamme phrygienne plus (low mumble) la vidéo sur le chiffrage de la partition.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 2.8/10; 11.2s, FR.
FR_BCoQVWd-9tk_W000025 · in -18.3 dBFS · gain -1.7 dB · emolia-02638
This chain comes from the one-sided rule: only Sourness had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Sourness barely there — 0.09, lower than 91 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.88.
Nothing was asked of the other axis, and in fact Affection drifts down from 0.93 to 0.03 (-0.90), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.25, then +0.24, then +0.16 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.43 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.39 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.43, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 34 s · bg · podcast
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.463 before conversion and 0.830 after — it rose by 0.368. Neighbour-to-neighbour the worst pair went 0.402 → 0.775. (The earlier render, with segment 1 left raw, scores 0.754 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sourness moved +0.879 in the original and +0.839 after conversion — 95 % of the delta retained, which is essentially all of it. On the other named axis, Affection, -0.687 became -0.009.
Quality. Mean predicted overall quality across the segments went 2.93 → 3.07 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.463 → 0.830+0.368identity cos neighbours 0.402 → 0.775d_b rescored +0.879 → +0.839d_a rescored -0.687 → -0.009d_a mined -0.896d_b mined 0.879min_cos_consec (site) 0.3901min_cos_anchor (site) 0.4270dataset podcastlang bgspeaker 118716total 33.2schain gain +1.0 dBseam step 3.5 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, balanced body, normally alert
(affection · fast, slightly relaxed, fairly steady, storytelling)странни японски неща, не се притеснявайте, целовки. Това е последното
full caption & clip details
A young adult feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as affection; style: storytelling, conversational; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 2.0/10; 4.5s, BG.
118716_00129304 · in -26.4 dBFS · gain +6.4 dB · podcast-02872
(affection, infatuation, contemplation·normal-paced, relaxed, fairly steady, casual)съобщение към родителите. Момичето, което си комуникира постоянно с родителите. Не комуникира повече с родители.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as affection, infatuation, contemplation; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.9/6; vocal-burst blend 6.1/10; 6.1s, BG.
118716_00129759 · in -25.2 dBFS · gain +5.2 dB · podcast-03834
(contemplation, shame·measured, relaxed, fairly steady, casual)в която живее. Ми, то мисля, че е масова практика да се изчаква някой, който не се появява до определен час да се изчака дали ще се появи и тогава, защото иначе той може да си пови се два часа и ти да си анжира на празно полицията.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, slightly guarded; reads as contemplation, shame; style: casual, whispered; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 9.2/10; 14.3s, BG.
118716_00132088 · in -27.3 dBFS · gain +7.3 dB · podcast-03852
(helplessness, distress, fatigue exhaustion· measured, slightly relaxed, fairly steady, casual)Еми да, но в случая те знаят, че тя е била на странно място с странен
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as helplessness, distress, fatigue exhaustion; style: casual, storytelling; average recording, no background noise; genuineness 3.2/6; vocal-burst blend 5.1/10; 4.3s, BG.
118716_00133544 · in -25.7 dBFS · gain +5.7 dB · podcast-02871
(sourness· measured, slightly relaxed, moderately variable, conversational)Държава се е супер странно, не иска да правим уроките, където се правят
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as sourness; style: conversational, casual; good recording, no background noise; genuineness 2.4/6; vocal-burst blend 3.3/10; 4.7s, BG.
118716_00135512 · in -24.9 dBFS · gain +4.9 dB · podcast-02865
This chain comes from the one-sided rule: only Fear had to get where it was going, by at least 0.70. The other emotion was left completely free.
The chain starts with Fear barely there — 0.09, lower than 91 % of clips in this corpus — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.79.
Nothing was asked of the other axis, and in fact Concentration drifts down from 0.97 to 0.71 (-0.26), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.18, then +0.20, then +0.24, then +0.16 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.85 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.85 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 51 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.800 before conversion and 0.825 after — it rose by 0.025. Neighbour-to-neighbour the worst pair went 0.922 → 0.858. (The earlier render, with segment 1 left raw, scores 0.435 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.786 in the original and +0.515 after conversion — 65 % of the delta retained. On the other named axis, Concentration, -0.256 became -0.212.
Quality. Mean predicted overall quality across the segments went 2.99 → 3.12 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.800 → 0.825+0.025identity cos neighbours 0.922 → 0.858d_b rescored +0.786 → +0.515d_a rescored -0.256 → -0.212d_a mined -0.256d_b mined 0.786min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_-Kaw3u-xVr0total 50.1schain gain +0.9 dBseam step 0.5 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(concentration · steady, newsreading, formal)In Callahan's view, Churchill was guilty of "'carefully reconstructing the story' to suit his post-war political goals.John Keegan wrote in the 1985 introduction to the series that some deficiencies in the account stem from the secrecy of ultra-intelligence.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 16.3s, EN.
EN_-Kaw3u-xVr0_W000029 · in -15.1 dBFS · gain -4.9 dB · emolia-00448
(fairly steady, newsreading, formal)Keegan held that Churchill's account was unique, since none of the other leaders – Franklin D. Roosevelt, Harry S. Truman, Benito Mussolini, Joseph Stalin, Adolf Hitler – wrote a first-hand account of the war
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 12.3s, EN.
EN_-Kaw3u-xVr0_W000030 · in -14.9 dBFS · gain -5.1 dB · emolia-00448
(malevolence malice·steady, formal, authoritative)Churchill's books were written collaboratively, as he solicited others involved in the war for their papers and remembrances
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice; style: formal, authoritative; good recording, no background noise; mildly explicit content; genuineness 0.2/6; vocal-burst blend 0.0/10; 6.6s, EN.
EN_-Kaw3u-xVr0_W000031 · in -14.3 dBFS · gain -5.7 dB · emolia-00448
(steady, newsreading, formal)The Second World War has been issued in editions of six, twelve and four volumes, as well as a single volume abridgment
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.2s, EN.
EN_-Kaw3u-xVr0_W000032 · in -14.4 dBFS · gain -5.6 dB · emolia-00448
(steady, formal, newsreading)Some volumes in these editions share names, such as triumph and tragedy but the contents of the volumes differ, covering varying portions of the book
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 8.5s, EN.
EN_-Kaw3u-xVr0_W000033 · in -14.4 dBFS · gain -5.6 dB · emolia-00448