c-emolia-AB2 — voice-corrected

Corpus emolia in isolation, rule AB2.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_c-emolia-AB2.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
53segments re-voiced
0.787 → 0.749median worst-to-anchor identity cosine
66 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Impatience and Irritability ↓  /  Jealousy and Envyidentity +0.23 emotion 56 %   c-emolia-AB2 · #1

This chain comes from the two-sided rule: it only counts if both emotions move — Impatience and Irritability down and Jealousy and Envy up — by at least 0.25 each.

The chain starts with Jealousy and Envy clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.30.

At the same time Impatience and Irritability goes the other way, from 0.87 (higher than 87 % of clips in this corpus) to 0.58 (higher than 58 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.06 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.55 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.55 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.55, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 33 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.495 before conversion and 0.720 after — it rose by 0.225. Neighbour-to-neighbour the worst pair went 0.495 → 0.720. (The earlier render, with segment 1 left raw, scores 0.622 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Jealousy and Envy moved +0.301 in the original and +0.170 after conversion — 56 % of the delta retained. On the other named axis, Impatience and Irritability, -0.291 became -0.345.

Quality. Mean predicted overall quality across the segments went 2.77 → 2.89 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.495 → 0.720 +0.225identity cos neighbours 0.495 → 0.720d_b rescored +0.301 → +0.170d_a rescored -0.291 → -0.345d_a mined -0.292d_b mined 0.301min_cos_consec (site) 0.5455min_cos_anchor (site) 0.5455dataset emolialang enspeaker EN_p2b-2zBJZUQtotal 32.4schain gain +1.8 dBseam step 1.1 dBcrossfades 100/100 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, quiet background, moderately variable, some disfluency, light breath
(normal-paced, normally alert, slightly relaxed, dramatic) The project is not the picture in the end, it's actually meeting the women.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: dramatic, storytelling; good recording, quiet background; genuineness 1.8/6; vocal-burst blend 0.8/10; 5.0s, EN.
EN_p2b-2zBJZUQ_W000016 · in -16.4 dBFS · gain -3.6 dB · emolia-02345
(affection, disappointment, contentment · normal-paced, normally alert, neutral tension, casual) That's the project. And it's all about our energy, like sometimes we (ahem) had to cancel a shoot, like we were in France and we were so tired so we just called the, the woman and say, we are so sorry but we would not be available, as available as we want to share this moment with you.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as affection, disappointment, contentment; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 6.7/10; 17.9s, EN.
EN_p2b-2zBJZUQ_W000017 · in -21.4 dBFS · gain +1.4 dB · emolia-02345
(jealousy and envy, contentment · brisk, energised, neutral tension, casual) It's always like in (low mumble) an environment that the woman chose. (ahem) And we ask the model usually like, do you want to be at your place? Do you want to be outdoor? Do you like water? Do you...
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as jealousy and envy, contentment; style: casual, playful; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 6.9/10; 9.9s, EN.
EN_p2b-2zBJZUQ_W000018 · in -15.7 dBFS · gain -4.3 dB · emolia-02345
Concentration ↓  /  Emotional Numbnessidentity +0.13 emotion 49 %   c-emolia-AB2 · #2

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.29.

At the same time Concentration goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.35. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.16, then +0.11, then -0.11, then +0.12 — not a clean run: step 3 moves back the other way by 0.11 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.67 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.68 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.67, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 42 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.518 before conversion and 0.643 after — it rose by 0.125. Neighbour-to-neighbour the worst pair went 0.681 → 0.790. (The earlier render, with segment 1 left raw, scores 0.529 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.290 in the original and +0.143 after conversion — 49 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Concentration, -0.354 became -0.303.

Quality. Mean predicted overall quality across the segments went 2.80 → 2.98 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.518 → 0.643 +0.125identity cos neighbours 0.681 → 0.790d_b rescored +0.290 → +0.143d_a rescored -0.354 → -0.303d_a mined -0.354d_b mined 0.291min_cos_consec (site) 0.6751min_cos_anchor (site) 0.6715dataset emolialang enspeaker EN_Elu3YMNSAZototal 40.9schain gain +2.3 dBseam step 0.8 dBcrossfades 150/100/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(concentration, contemplation · steady, almost no disfluency, minimal breath, newsreading) The test is intended to draw the line between legitimate criticism towards the State of Israel, its actions and policies, and non-legitimate criticism that becomes anti-Semitic. Earl Raab writes that "[t], here is a new surge of anti-Semitism in the world, and much prejudice against Israel is driven by such anti-Semitism."
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, contemplation; style: newsreading, authoritative; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 19.6s, EN.
EN_Elu3YMNSAZo_W000079 · in -16.3 dBFS · gain -3.7 dB · emolia-00448
(contempt · fairly steady, no disfluency, light breath, formal) But argues that charges of antisemitism based on anti-Israel opinions generally lack credibility.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt; style: formal, authoritative; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 1.1/10; 6.0s, EN.
EN_Elu3YMNSAZo_W000080 · in -16.1 dBFS · gain -3.9 dB · emolia-00448
(emotional numbness · steady, no disfluency, light breath, formal) This reduces the problems of prejudice against Israel to cartoon proportions.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, authoritative; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.7/10; 5.2s, EN.
EN_Elu3YMNSAZo_W000081 · in -14.9 dBFS · gain -5.0 dB · emolia-00448
(contempt, disgust · steady, almost no disfluency, light breath, formal) Raab describes prejudice against Israel as a "...serious breach of morality and good sense
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, disgust; style: formal, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.0/10; 6.2s, EN.
EN_Elu3YMNSAZo_W000082 · in -15.4 dBFS · gain -4.6 dB · emolia-00448
(emotional numbness · steady, no disfluency, light breath, formal) Part of what a reasonably informed, progressive, decent person thinks
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.5/10; 4.5s, EN.
EN_Elu3YMNSAZo_W000083 · in -14.8 dBFS · gain -5.2 dB · emolia-00448
Emotional Numbness ↓  /  Concentrationidentity −0.06 emotion 83 %   c-emolia-AB2 · #3

This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Concentration up — by at least 0.25 each.

The chain starts with Concentration clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.37.

At the same time Emotional Numbness goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.59 (higher than 59 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.18, then +0.08, then +0.11 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 50 s · fr · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.870 before conversion and 0.809 after — it fell by 0.060. Neighbour-to-neighbour the worst pair went 0.868 → 0.845. (The earlier render, with segment 1 left raw, scores 0.761 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.366 in the original and +0.303 after conversion — 83 % of the delta retained, which is most of it. On the other named axis, Emotional Numbness, -0.402 became -0.319.

Quality. Mean predicted overall quality across the segments went 3.12 → 3.21 (+0.09) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.870 → 0.809 -0.060identity cos neighbours 0.868 → 0.845d_b rescored +0.366 → +0.303d_a rescored -0.402 → -0.319d_a mined -0.402d_b mined 0.368min_cos_consec (site) 0.9081min_cos_anchor (site) 0.9081dataset emolialang frspeaker FR_E28QGC6bEvItotal 49.4schain gain +0.3 dBseam step 1.0 dBcrossfades 150/100/100 ms
Script — 4 chunks, 4 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, balanced body, average recording, quiet background, fairly steady, moderate pitch range, light breath
(emotional numbness, doubt, shame · normal-paced, normally alert, slightly relaxed, monologue) (ahem) euh, qu'on faisait de la compétence sans s'en rendre compte. C'est-à-dire que, (ahem) euh, y a pas tellement d'innovation.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, doubt, shame; style: monologue, didactic; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.0/10; 5.4s, FR.
FR_E28QGC6bEvI_W000054 · in -16.0 dBFS · gain -4.0 dB · emolia-02658
(distress, disappointment · normal-paced, normally alert, slightly relaxed, monologue) (ahem) En fait, on a fait du regroupement de choses que l'on faisait déjà. On s'est aperçu qu'effectivement,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as distress, disappointment; style: monologue, didactic; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.0/10; 8.3s, FR.
FR_E28QGC6bEvI_W000055 · in -18.5 dBFS · gain -1.5 dB · emolia-02658
(disappointment, anger, impatience and irritability · normal-paced, normally alert, slightly relaxed, monologue) On mettait déjà de manière, (low mumble) euh, c'est ce travail qui nous a permis de le, de le rendre plus clair en quelque sorte, mais on s'est bien aperçu qu'on mettait déjà en situation clairement les étudiants et que d'une certaine manière, il y avait plutôt un travail de clarification à faire et de rationalisation de ce qui était déjà fait. Donc, on a introduit peu de nouveautés.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as disappointment, anger, impatience and irritability; style: monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.8/10; 18.3s, FR.
FR_E28QGC6bEvI_W000056 · in -16.8 dBFS · gain -3.2 dB · emolia-02658
(concentration, contemplation, relief · measured, subdued, neutral tension, monologue) dans ces évaluations et ces mises en situation. Mais on a procédé essentiellement à des regroupements. Alors, peut-être que là, (low mumble) euh, ce qu'il y a de différent avec ce que j'ai vu tout à l'heure, euh, (low mumble) par rapport à la licence de STAPS, c'est qu'on a fait le choix, euh, (ahem) pour nous réalistes, (low mumble) euh, de regrouper
full caption & clip details
An adult masculine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, contemplation, relief; style: monologue, casual; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 1.3/10; 17.9s, FR.
FR_E28QGC6bEvI_W000057 · in -16.1 dBFS · gain -4.0 dB · emolia-02658
Sexual Lust ↓  /  Triumphidentity −0.05 emotion 64 %   c-emolia-AB2 · #4

This chain comes from the two-sided rule: it only counts if both emotions move — Sexual Lust down and Triumph up — by at least 0.25 each.

The chain starts with Triumph around average — 0.51, higher than 51 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.32.

At the same time Sexual Lust goes the other way, from 0.77 (higher than 77 % of clips in this corpus) to 0.50 (right about the corpus median), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.08 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.98 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.98 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.98), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 40 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.963 before conversion and 0.917 after — it fell by 0.046. Neighbour-to-neighbour the worst pair went 0.963 → 0.917. (The earlier render, with segment 1 left raw, scores 0.803 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Triumph moved +0.324 in the original and +0.208 after conversion — 64 % of the delta retained. On the other named axis, Sexual Lust, -0.258 became -0.359.

Quality. Mean predicted overall quality across the segments went 2.94 → 3.20 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.963 → 0.917 -0.046identity cos neighbours 0.963 → 0.917d_b rescored +0.324 → +0.208d_a rescored -0.258 → -0.359d_a mined -0.268d_b mined 0.323min_cos_consec (site) 0.9818min_cos_anchor (site) 0.9815dataset emolialang enspeaker EN_DByRrF9akYItotal 38.9schain gain +2.3 dBseam step 0.6 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, good recording, no background noise, normal-paced, normally alert, slightly relaxed, clear
(fairly steady, no disfluency, formal, newsreading) In 1993, the Wadia Group acquired a stake in Associated Biscuits International, ABIL, and became an equal partner with Group Dannon in Britannia Industries Limited
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 9.9s, EN.
EN_DByRrF9akYI_W000012 · in -17.5 dBFS · gain -2.5 dB · emolia-00609
(sadness · steady, almost no disfluency, newsreading, formal) In what the Economic Times referred to as one of ''India's'', most dramatic corporate sagas, Palai ceded control to Wadia and Danon after a bitter boardroom struggle, then fled his Singapore base to India in 1995 after accusations of defrauding Britannia, and died the same year in Tihar jail.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sadness; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 16.2s, EN.
EN_DByRrF9akYI_W000013 · in -17.7 dBFS · gain -2.3 dB · emolia-00609
(fairly steady, no disfluency, newsreading, formal) The Wadears-Kalabakan Investments and Group Dannon had two equal joint venture companies, Wadia BSN and United Kingdom registered Associated Biscuits International Holdings Limited, which together held a 51% stake in Britannia
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 13.2s, EN.
EN_DByRrF9akYI_W000014 · in -13.9 dBFS · gain -6.0 dB · emolia-00609
Concentration ↓  /  Fearidentity −0.09 emotion 110 %   c-emolia-AB2 · #5

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Fear up — by at least 0.25 each.

The chain starts with Fear clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.27.

At the same time Concentration goes the other way, from 0.87 (higher than 87 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.07 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 23 s · fr · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.856 before conversion and 0.764 after — it fell by 0.093. Neighbour-to-neighbour the worst pair went 0.856 → 0.764. (The earlier render, with segment 1 left raw, scores 0.770 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.271 in the original and +0.297 after conversion — 110 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.259 became -0.151.

Quality. Mean predicted overall quality across the segments went 2.71 → 3.07 (+0.36) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.856 → 0.764 -0.093identity cos neighbours 0.856 → 0.764d_b rescored +0.271 → +0.297d_a rescored -0.259 → -0.151d_a mined -0.260d_b mined 0.271min_cos_consec (site) 0.8748min_cos_anchor (site) 0.8748dataset emolialang frspeaker FR_lPdeTeb5EGktotal 22.4schain gain +1.7 dBseam step 0.9 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, normally alert, slightly relaxed, light breath
(measured, fairly steady, some disfluency, monologue) pire encore, les productions maraîchères basées sur la monoculture et l'utilisation d'intrants chimiques commencent à être remises en cause par les producteurs humains.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 5.3/10; 9.0s, FR.
FR_lPdeTeb5EGk_W000002 · in -15.8 dBFS · gain -4.2 dB · emolia-02820
(sadness · normal-paced, steady, no disfluency, narration) Si l'agriculture conventionnelle pose d'importants problèmes de santé, ce n'est pas le seul défaut de ce système.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sadness; style: narration, formal; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.2/10; 5.2s, FR.
FR_lPdeTeb5EGk_W000003 · in -16.1 dBFS · gain -3.9 dB · emolia-02820
(fear · measured, steady, almost no disfluency, monologue) Pour pallier à toutes ces difficultés, quelques producteurs se sont tournés vers l'agriculture agroécologique. Un mouvement qui tend à se développer en Côte d'Ivoire.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear; style: monologue, narration; average recording, quiet background; genuineness 0.8/6; vocal-burst blend 2.3/10; 8.6s, FR.
FR_lPdeTeb5EGk_W000004 · in -17.6 dBFS · gain -2.4 dB · emolia-02820
Pain ↓  /  Emotional Numbnessidentity −0.03 emotion 7 %   c-emolia-AB2 · #6

This chain comes from the two-sided rule: it only counts if both emotions move — Pain down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.25.

At the same time Pain goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.35 (lower than 65 % of clips in this corpus), a change of -0.51. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.23, then -0.02, then -0.10, then +0.15 — not a clean run: step 2 moves back the other way by 0.02 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 30 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.846 before conversion and 0.817 after — it fell by 0.030. Neighbour-to-neighbour the worst pair went 0.893 → 0.814. (The earlier render, with segment 1 left raw, scores 0.778 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.252 in the original and +0.018 after conversion — 7 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Pain, -0.514 became +0.348.

Quality. Mean predicted overall quality across the segments went 2.86 → 2.95 (+0.09) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.846 → 0.817 -0.030identity cos neighbours 0.893 → 0.814d_b rescored +0.252 → +0.018d_a rescored -0.514 → +0.348d_a mined -0.514d_b mined 0.252min_cos_consec (site) 0.8900min_cos_anchor (site) 0.8471dataset emolialang enspeaker EN_B00058_S00272total 29.1schain gain +1.2 dBseam step 0.7 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(normal-paced, no disfluency, formal, monologue) To shorten or lengthen the transition in the timeline, with it selected, drag either trim handle.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 2.3/10; 5.0s, EN.
EN_B00058_S00272_W000005 · in -18.4 dBFS · gain -1.6 dB · emolia-01357
(emotional numbness · measured, no disfluency, formal, monologue) If you want to replace the transition type drag another down and release it on top to perform a replace edit. If you want to delete a transition
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.0/10; 7.5s, EN.
EN_B00058_S00272_W000006 · in -20.1 dBFS · gain +0.1 dB · emolia-01357
(normal-paced, no disfluency, formal, narration) Either tap the trash can or drag the transition off the timeline and release it over the viewer.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 1.7/10; 4.6s, EN.
EN_B00058_S00272_W000007 · in -19.5 dBFS · gain -0.5 dB · emolia-01357
(jealousy and envy · normal-paced, little disfluency, formal, narration) The wipe and slide transitions are really useful for title reveals. I have two title tracks over my main video track.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as jealousy and envy; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.3/10; 6.8s, EN.
EN_B00058_S00272_W000008 · in -19.3 dBFS · gain -0.7 dB · emolia-01357
(emotional numbness · measured, no disfluency, formal, narration) And I'll drag a wipe down transition onto one and a slide up transition on the other and preview the results.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, narration; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 1.0/10; 6.0s, EN.
EN_B00058_S00272_W000009 · in -19.5 dBFS · gain -0.5 dB · emolia-01357
Contemplation ↓  /  Painidentity +0.02 emotion REVERSED   c-emolia-AB2 · #7

This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Pain up — by at least 0.25 each.

The chain starts with Pain around average — 0.54, higher than 54 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.31.

At the same time Contemplation goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.62 (higher than 62 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.13 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 19 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.854 before conversion and 0.871 after — it rose by 0.017. Neighbour-to-neighbour the worst pair went 0.869 → 0.859. (The earlier render, with segment 1 left raw, scores 0.824 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The emotional move did not survive. Re-scored end to end, Pain moved +0.305 in the original and -0.196 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Contemplation, -0.312 became -0.372.

Quality. Mean predicted overall quality across the segments went 2.86 → 3.17 (+0.31) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.854 → 0.871 +0.017identity cos neighbours 0.869 → 0.859d_b rescored +0.305 → -0.196d_a rescored -0.312 → -0.372d_a mined -0.312d_b mined 0.305min_cos_consec (site) 0.8775min_cos_anchor (site) 0.8887dataset emolialang zhspeaker ZH_B00005_S04669total 18.2schain gain +1.9 dBseam step 0.9 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady, moderate pitch range
(contemplation · measured, some disfluency, clear, monologue) 米开朗基罗失手掉了一些工具下来差点砸在教皇头上,教皇非常生气。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation; style: monologue, whispered; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 3.7/10; 8.0s, ZH.
ZH_B00005_S04669_W000023 · in -17.7 dBFS · gain -2.3 dB · emolia-03331
(measured, no disfluency, clear, formal) 世界各地的人们都赶来参观这个天花板,可是总仰着头看很不方便。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, whispered; average recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.7/10; 6.8s, ZH.
ZH_B00005_S04669_W000024 · in -17.6 dBFS · gain -2.4 dB · emolia-03331
(normal-paced, some disfluency, average clarity, monologue) 想要舒舒服服的看,就只能躺在地板上。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 1.6/10; 3.8s, ZH.
ZH_B00005_S04669_W000025 · in -21.8 dBFS · gain +1.8 dB · emolia-03331
Concentration ↓  /  Fearidentity −0.10 emotion 157 %   c-emolia-AB2 · #8

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Fear up — by at least 0.25 each.

The chain starts with Fear around average — 0.54, higher than 54 % of clips in this corpus — and ends with it strongly present at 0.81, higher than 81 % of clips in this corpus. That is a total rise of 0.27.

At the same time Concentration goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.08, then +0.19 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 34 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.825 before conversion and 0.720 after — it fell by 0.104. Neighbour-to-neighbour the worst pair went 0.872 → 0.764. (The earlier render, with segment 1 left raw, scores 0.662 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.273 in the original and +0.428 after conversion — 157 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.283 became -0.251.

Quality. Mean predicted overall quality across the segments went 2.90 → 3.10 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.825 → 0.720 -0.104identity cos neighbours 0.872 → 0.764d_b rescored +0.273 → +0.428d_a rescored -0.283 → -0.251d_a mined -0.283d_b mined 0.273min_cos_consec (site) 0.9008min_cos_anchor (site) 0.9252dataset emolialang enspeaker EN_B00039_S04481total 32.9schain gain +3.1 dBseam step 0.7 dBcrossfades 150/100 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, normally alert, slightly relaxed, average clarity
(concentration, interest, intoxication altered states of consciousness · measured, moderately variable, frequent disfluency, conversational) On the high level, I think you can use, uh, (ahem) these data-driven, (low mumble) uh, composition tools for a variety of different methods. So you can, you know, use things that are 3D printed, things that are non-3D printed. So this is, I don't think that, I don't sort of see this as a, as a, as a big problem when you actually build, (low mumble) uh,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as concentration, interest, intoxication altered states of consciousness; style: conversational, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 4.7/10; 19.0s, EN.
EN_B00039_S04481_W000218 · in -19.4 dBFS · gain -0.6 dB · emolia-00993
(interest, concentration · normal-paced, fairly steady, frequent disfluency, casual) Uh, (low mumble) things from components. And there's a large category of, of objects that actually are build from components, you know.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, concentration; style: casual, monologue; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 0.9/10; 7.8s, EN.
EN_B00039_S04481_W000219 · in -19.8 dBFS · gain -0.2 dB · emolia-00993
(normal-paced, fairly steady, some disfluency, casual) There are a number of industrial (low mumble) machines, both 3D systems and Stratasys have these machines. As I said,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.2/10; 6.4s, EN.
EN_B00039_S04481_W000220 · in -18.2 dBFS · gain -1.8 dB · emolia-00993
Fatigue Exhaustion ↓  /  Triumphidentity +0.03 emotion REVERSED   c-emolia-AB2 · #9

This chain comes from the two-sided rule: it only counts if both emotions move — Fatigue Exhaustion down and Triumph up — by at least 0.25 each.

The chain starts with Triumph around average — 0.51, higher than 51 % of clips in this corpus — and ends with it strongly present at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.28.

At the same time Fatigue Exhaustion goes the other way, from 0.75 (higher than 75 % of clips in this corpus) to 0.35 (lower than 65 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.05, then +0.23 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.77 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.77 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.77, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 18 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.754 before conversion and 0.782 after — it rose by 0.028. Neighbour-to-neighbour the worst pair went 0.754 → 0.790. (The earlier render, with segment 1 left raw, scores 0.707 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Triumph moved +0.285 in the original and -0.178 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Fatigue Exhaustion, -0.395 became -0.466.

Quality. Mean predicted overall quality across the segments went 2.63 → 2.86 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.754 → 0.782 +0.028identity cos neighbours 0.754 → 0.790d_b rescored +0.285 → -0.178d_a rescored -0.395 → -0.466d_a mined -0.395d_b mined 0.285min_cos_consec (site) 0.7691min_cos_anchor (site) 0.7691dataset emolialang enspeaker EN_B00044_S06796total 17.1schain gain +1.2 dBseam step 0.7 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · slightly cool, neutral-bright, fairly smooth, balanced body, normal-paced, energised, slightly relaxed, moderate pitch range
(fairly steady, almost no disfluency, clear, authoritative) This role requires communicating and coordinating with vendors as well.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.4/10; 4.8s, EN.
EN_B00044_S06796_W000006 · in -13.3 dBFS · gain -6.7 dB · emolia-01100
(fairly steady, some disfluency, very clear, authoritative) So these Azure administrators use the Azure portal as they become more proficient, they use
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; very clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, dramatic; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 0.2/10; 7.0s, EN.
EN_B00044_S06796_W000007 · in -13.9 dBFS · gain -6.1 dB · emolia-01100
(moderately variable, some disfluency, average clarity, playful) Azure administrators mainly use the Azure portal and as they become more provisioned,
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: playful, storytelling; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 1.4/10; 5.6s, EN.
EN_B00044_S06796_W000008 · in -13.8 dBFS · gain -6.2 dB · emolia-01100
Disgust ↓  /  Contemplationidentity +0.05 emotion 69 %   c-emolia-AB2 · #10

This chain comes from the two-sided rule: it only counts if both emotions move — Disgust down and Contemplation up — by at least 0.25 each.

The chain starts with Contemplation clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.34.

At the same time Disgust goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.15 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 27 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.807 before conversion and 0.862 after — it rose by 0.054. Neighbour-to-neighbour the worst pair went 0.807 → 0.862. (The earlier render, with segment 1 left raw, scores 0.856 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.338 in the original and +0.232 after conversion — 69 % of the delta retained. On the other named axis, Disgust, -0.307 became -0.833.

Quality. Mean predicted overall quality across the segments went 2.78 → 3.13 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.807 → 0.862 +0.054identity cos neighbours 0.807 → 0.862d_b rescored +0.338 → +0.232d_a rescored -0.307 → -0.833d_a mined -0.307d_b mined 0.339min_cos_consec (site) 0.8475min_cos_anchor (site) 0.8475dataset emolialang enspeaker EN_voem3Nfrogytotal 26.7schain gain +2.6 dBseam step 0.6 dBcrossfades 100/100 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned
(disgust · normal-paced, normally alert, fully relaxed, casual) It's also called young people. (ahem) Uhm, it's, it's real world alleged (low mumble) result is
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as disgust; style: casual, conversational; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 1.2/10; 7.1s, EN.
EN_voem3Nfrogy_W000091 · in -17.8 dBFS · gain -2.2 dB · emolia-02507
(emotional numbness, fear, sadness · slow, very low-energy, slightly relaxed, whispered) The lack of, of demand for insurance. Lots of, from a neoclassical point of view, (ahem) uh, people seem to under insure. They do not.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness, fear, sadness; style: whispered, monologue; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 0.6/10; 11.6s, EN.
EN_voem3Nfrogy_W000092 · in -14.9 dBFS · gain -5.1 dB · emolia-02507
(contemplation, sexual lust, concentration · measured, normally alert, slightly relaxed, monologue) Buy products that exist that would (ahem) allow them to buy the mean of an outcome and reduce the variance that might occur.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as contemplation, sexual lust, concentration; style: monologue, whispered; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.7/10; 8.3s, EN.
EN_voem3Nfrogy_W000093 · in -15.0 dBFS · gain -5.0 dB · emolia-02507
Infatuation ↓  /  Emotional Numbnessidentity +0.01 emotion 44 %   c-emolia-AB2 · #11

This chain comes from the two-sided rule: it only counts if both emotions move — Infatuation down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness around average — 0.44, lower than 56 % of clips in this corpus — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.39.

At the same time Infatuation goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.39 (lower than 61 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.18 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.74 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.70 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.74, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 19 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.691 before conversion and 0.706 after — it rose by 0.015. Neighbour-to-neighbour the worst pair went 0.695 → 0.735. (The earlier render, with segment 1 left raw, scores 0.636 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.387 in the original and +0.172 after conversion — 44 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Infatuation, -0.328 became -0.121.

Quality. Mean predicted overall quality across the segments went 2.59 → 2.83 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.691 → 0.706 +0.015identity cos neighbours 0.695 → 0.735d_b rescored +0.387 → +0.172d_a rescored -0.328 → -0.121d_a mined -0.328d_b mined 0.387min_cos_consec (site) 0.6967min_cos_anchor (site) 0.7387dataset emolialang enspeaker EN_VXkPIxzzWuEtotal 18.2schain gain +3.0 dBseam step 1.7 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, normally alert, slightly relaxed, fairly steady
(measured, frequent disfluency, somewhat unclear, casual) 38 lands, a lot of these, (low mumble) uh, so we've got 11 islands, 15 swamps, a lot of lands that are just blue-black lands, so yeaah.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 1.1/10; 9.3s, EN.
EN_VXkPIxzzWuE_W000045 · in -18.1 dBFS · gain -1.9 dB · emolia-01177
(fatigue exhaustion · measured, some disfluency, slurred, casual) So, estuary, hang on, what I'll do is I'll change it to visual spoiler.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as fatigue exhaustion; style: casual, monologue; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 2.5/10; 4.4s, EN.
EN_VXkPIxzzWuE_W000046 · in -18.2 dBFS · gain -1.8 dB · emolia-01177
(normal-paced, some disfluency, slurred, casual) This will probably take a while to light up. Command tower, of course, pretty standard in commander.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 1.9/10; 4.7s, EN.
EN_VXkPIxzzWuE_W000047 · in -19.0 dBFS · gain -1.0 dB · emolia-01177
Confusion ↓  /  Fatigue Exhaustionidentity −0.03 emotion 129 %   c-emolia-AB2 · #12

This chain comes from the two-sided rule: it only counts if both emotions move — Confusion down and Fatigue Exhaustion up — by at least 0.25 each.

The chain starts with Fatigue Exhaustion around average — 0.57, higher than 57 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.32.

At the same time Confusion goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.10, then +0.23 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 20 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.760 before conversion and 0.734 after — it fell by 0.027. Neighbour-to-neighbour the worst pair went 0.760 → 0.734. (The earlier render, with segment 1 left raw, scores 0.687 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.324 in the original and +0.418 after conversion — 129 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Confusion, -0.316 became +0.094.

Quality. Mean predicted overall quality across the segments went 2.89 → 2.97 (+0.08) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.760 → 0.734 -0.027identity cos neighbours 0.760 → 0.734d_b rescored +0.324 → +0.418d_a rescored -0.316 → +0.094d_a mined -0.316d_b mined 0.324min_cos_consec (site) 0.8312min_cos_anchor (site) 0.8473dataset emolialang zhspeaker ZH_B00008_S06791total 19.8schain gain +2.1 dBseam step 1.9 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(confusion · measured, some disfluency, somewhat unclear, monologue) 问我使用止损指令,因为我无法持续的观察自己的股票价格做市商。为了挣出止损指令,情况会有多么普遍。我的头寸大约是十万美元。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as confusion; style: monologue, didactic; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 5.6/10; 10.0s, ZH.
ZH_B00008_S06791_W000032 · in -18.6 dBFS · gain -1.4 dB · emolia-03353
(normal-paced, no disfluency, clear, formal) 要是你持有的股票出现跳空下跌缺口会怎么办?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, conversational; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 5.7/10; 3.2s, ZH.
ZH_B00008_S06791_W000033 · in -20.6 dBFS · gain +0.6 dB · emolia-03353
(measured, some disfluency, somewhat unclear, monologue) 通常情况下,大多数这种缺口都应该卖出。如果该股票啊,我理解跟红旗可能是个警报啊。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, authoritative; average recording, no background noise; genuineness 3.0/6; vocal-burst blend 4.8/10; 6.9s, ZH.
ZH_B00008_S06791_W000034 · in -19.4 dBFS · gain -0.6 dB · emolia-03353
Sourness ↓  /  Concentrationidentity +0.51 emotion 61 %   c-emolia-AB2 · #13

This chain comes from the two-sided rule: it only counts if both emotions move — Sourness down and Concentration up — by at least 0.25 each.

The chain starts with Concentration clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.27.

At the same time Sourness goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.05, then +0.22 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.05 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.05 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.05, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 27 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.048 before conversion and 0.466 after — it rose by 0.513. Neighbour-to-neighbour the worst pair went 0.051 → 0.447. (The earlier render, with segment 1 left raw, scores 0.382 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.270 in the original and +0.164 after conversion — 61 % of the delta retained. On the other named axis, Sourness, -0.289 became -0.438.

Quality. Mean predicted overall quality across the segments went 2.89 → 3.01 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 -0.048 → 0.466 +0.513identity cos neighbours 0.051 → 0.447d_b rescored +0.270 → +0.164d_a rescored -0.289 → -0.438d_a mined -0.289d_b mined 0.271min_cos_consec (site) 0.0453min_cos_anchor (site) -0.0475dataset emolialang enspeaker EN_PQh0ab6JRHYtotal 26.5schain gain +4.5 dBseam step 1.1 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, slightly relaxed, fairly steady, average clarity, moderate pitch range, light breath
(sourness · brisk, normally alert, little disfluency, casual) The analytics specialist in your organization, like a data scientist or a data engineer, all the way out to maybe a product manager or someone who doesn't really think of them as an analytics expert.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as sourness; style: casual, conversational; good recording, quiet background; genuineness 1.0/6; vocal-burst blend 4.5/10; 9.4s, EN.
EN_PQh0ab6JRHY_W000018 · in -20.8 dBFS · gain +0.8 dB · emolia-02161
(normal-paced, normally alert, little disfluency, casual) (low mumble) Uhm, using Alation, (low mumble) uhm, either directly or sometimes through one of our partnerships with folks like Tableau or MicroStrategy or Power BI.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 2.1/10; 8.1s, EN.
EN_PQh0ab6JRHY_W000019 · in -21.5 dBFS · gain +1.5 dB · emolia-02161
(concentration, contemplation · measured, subdued, frequent disfluency, monologue) So if we think about this notion of self-service analytics, Stephanie, and again, it's, (low mumble) uh, Elation has been a leader in defining this overall category.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as concentration, contemplation; style: monologue, casual; average recording, no background noise; genuineness 2.9/6; vocal-burst blend 0.0/10; 9.4s, EN.
EN_PQh0ab6JRHY_W000020 · in -26.4 dBFS · gain +6.4 dB · emolia-02161
Emotional Numbness ↓  /  Concentrationidentity −0.19 emotion 83 %   c-emolia-AB2 · #14

This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Concentration up — by at least 0.25 each.

The chain starts with Concentration around average — 0.56, higher than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.41.

At the same time Emotional Numbness goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.55 (higher than 55 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.21, then -0.05, then +0.06 — not a clean run: step 3 moves back the other way by 0.05 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 53 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.796 before conversion and 0.603 after — it fell by 0.193. Neighbour-to-neighbour the worst pair went 0.891 → 0.603. (The earlier render, with segment 1 left raw, scores 0.569 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.409 in the original and +0.338 after conversion — 83 % of the delta retained, which is most of it. On the other named axis, Emotional Numbness, -0.358 became -0.523.

Quality. Mean predicted overall quality across the segments went 2.96 → 3.08 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.796 → 0.603 -0.193identity cos neighbours 0.891 → 0.603d_b rescored +0.409 → +0.338d_a rescored -0.358 → -0.523d_a mined -0.358d_b mined 0.409min_cos_consec (site) 0.8788min_cos_anchor (site) 0.8735dataset emolialang enspeaker EN_dLfd2LY-WA0total 51.8schain gain +2.6 dBseam step 1.0 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, balanced body, average recording, quiet background, normally alert, slightly relaxed
(emotional numbness · normal-paced, fairly steady, frequent disfluency, casual) The only action that the facilitator could do when a user is, uh. (low mumble)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as emotional numbness; style: casual, monologue; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 0.6/10; 4.8s, EN.
EN_dLfd2LY-WA0_W000134 · in -18.9 dBFS · gain -1.1 dB · emolia-00625
(measured, steady, frequent disfluency, monologue) Too wrong, it's too off track. (ahem) He's tried to, okay, (chuckle) correct them if they really cannot recover.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 0.6/10; 7.0s, EN.
EN_dLfd2LY-WA0_W000135 · in -20.7 dBFS · gain +0.7 dB · emolia-00625
(doubt, confusion, concentration · normal-paced, fairly steady, some disfluency, casual) Because maybe it looks like, (ahem) I don't know, another attachment or looks like a decoration or (ahem) the label of the button is totally different from what the user is expecting. So after a while, if you see that the user
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, confusion, concentration; style: casual, monologue; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.8/10; 13.6s, EN.
EN_dLfd2LY-WA0_W000136 · in -20.6 dBFS · gain +0.6 dB · emolia-00625
(concentration · normal-paced, fairly steady, some disfluency, casual) Cannot (ahem) proceed, then you can give (ahem) a suggestion. Or if the user went in a totally wrong direction, then you can say, okay, this is not the right way, let's go back.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: casual, monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 1.6/10; 11.0s, EN.
EN_dLfd2LY-WA0_W000137 · in -19.9 dBFS · gain -0.1 dB · emolia-00625
(concentration, pain, disappointment · normal-paced, fairly steady, frequent disfluency, casual) Of course, if this happens, it's a failure. (low mumble) It's a big failure of the interface, (low mumble) and of course it's important to correct it, but also we should have the user proceed so that we may complete the test
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, pain, disappointment; style: casual, monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 2.0/10; 16.3s, EN.
EN_dLfd2LY-WA0_W000138 · in -21.1 dBFS · gain +1.1 dB · emolia-00625
Confusion ↓  /  Impatience and Irritabilityidentity +0.05 emotion 161 %   c-emolia-AB2 · #15

This chain comes from the two-sided rule: it only counts if both emotions move — Confusion down and Impatience and Irritability up — by at least 0.25 each.

The chain starts with Impatience and Irritability clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.29.

At the same time Confusion goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.04 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 32 s · ko · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.832 before conversion and 0.886 after — it rose by 0.054. Neighbour-to-neighbour the worst pair went 0.882 → 0.866. (The earlier render, with segment 1 left raw, scores 0.836 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.287 in the original and +0.463 after conversion — 161 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Confusion, -0.299 became -0.498.

Quality. Mean predicted overall quality across the segments went 2.91 → 3.20 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.832 → 0.886 +0.054identity cos neighbours 0.882 → 0.866d_b rescored +0.287 → +0.463d_a rescored -0.299 → -0.498d_a mined -0.299d_b mined 0.289min_cos_consec (site) 0.8721min_cos_anchor (site) 0.8274dataset emolialang kospeaker KO_GidUmZ9_P1gtotal 30.9schain gain +1.3 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(confusion, distress, longing · somewhat unclear, didactic, conversational) 같은 tutu로 묶으면 좀 장비 효율성이 좀 떨어지죠. 장비를 활용한 그, 관절기, 관절 동원 부분에서의 개념이 좀 동일한 개념이기 때문에.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, distress, longing; style: didactic, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 6.3/10; 10.4s, KO.
KO_GidUmZ9_P1g_W000050 · in -17.2 dBFS · gain -2.8 dB · emolia-03131
(distress, pain, jealousy and envy · average clarity, didactic, conversational) 그 다음에 자기가 머릿속으로, 아, 외상방 구조를 생각을 하셔야 돼. 그 다음에 내가 외상방으로 뻗는 게 어떻게 보면 가장 롱 스트레칭이 일어나는 거겠지. 그니까 그래서 여기까지 올라와줘야 된단 말이야. 그 다음에 풀 컨트렉션 들어가는 거죠. 시작.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as distress, pain, jealousy and envy; style: didactic, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 6.3/10; 13.1s, KO.
KO_GidUmZ9_P1g_W000051 · in -16.9 dBFS · gain -3.1 dB · emolia-03131
(impatience and irritability, fatigue exhaustion · average clarity, didactic, conversational) stop. 거기서 단축시키지 말고 바로 신전으로 바꿔서, 세, 아니지. 휙, 이러고 올라갔잖아. 통제, 컨트롤이 전혀 안 됐잖아. 시작.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, fatigue exhaustion; style: didactic, conversational; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 4.4/10; 7.8s, KO.
KO_GidUmZ9_P1g_W000052 · in -20.1 dBFS · gain +0.1 dB · emolia-03131
Astonishment Surprise ↓  /  Emotional Numbnessidentity +0.05 emotion 107 %   c-emolia-AB2 · #16

This chain comes from the two-sided rule: it only counts if both emotions move — Astonishment Surprise down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness around average — 0.55, higher than 55 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.34.

At the same time Astonishment Surprise goes the other way, from 0.79 (higher than 79 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.15 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.80 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.80. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 20 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.790 before conversion and 0.835 after — it rose by 0.045. Neighbour-to-neighbour the worst pair went 0.695 → 0.790. (The earlier render, with segment 1 left raw, scores 0.764 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.339 in the original and +0.364 after conversion — 107 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Astonishment Surprise, -0.266 became -0.257.

Quality. Mean predicted overall quality across the segments went 3.05 → 3.20 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.790 → 0.835 +0.045identity cos neighbours 0.695 → 0.790d_b rescored +0.339 → +0.364d_a rescored -0.266 → -0.257d_a mined -0.266d_b mined 0.340min_cos_consec (site) 0.8136min_cos_anchor (site) 0.8011dataset emolialang zhspeaker ZH_B00010_S06522total 19.5schain gain +1.8 dBseam step 0.7 dBcrossfades 100/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(measured, formal, monologue) 真正麻烦的事儿在于,好价格是不一定会出现的。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 2.6/10; 5.0s, ZH.
ZH_B00010_S06522_W000002 · in -18.4 dBFS · gain -1.6 dB · emolia-03374
(measured, didactic, monologue) 如果自己看懂的一家企业,能以很低的价格买入,然后之后公司能够大幅度的发展获得回报,这是最美妙的事儿。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 5.3/10; 10.6s, ZH.
ZH_B00010_S06522_W000003 · in -18.9 dBFS · gain -1.1 dB · emolia-03374
(normal-paced, formal, authoritative) 就像三公消费下的茅台一样,想想就流口水。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 2.8/10; 4.1s, ZH.
ZH_B00010_S06522_W000004 · in -17.9 dBFS · gain -2.1 dB · emolia-03374
Pride ↓  /  Hope Enthusiasm Optimismidentity +0.06 emotion 25 %   c-emolia-AB2 · #17

This chain comes from the two-sided rule: it only counts if both emotions move — Pride down and Hope Enthusiasm Optimism up — by at least 0.25 each.

The chain starts with Hope Enthusiasm Optimism strongly present — 0.88, higher than 88 % of clips in this corpus — and works its way down to clearly present at 0.62, higher than 62 % of clips in this corpus. That is a total fall of 0.25.

At the same time Pride goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.45. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.00, then -0.11, then -0.19, then +0.05 — not a clean run: step 1 moves back the other way by 0.00 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.60 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.73 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.60, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 22 s · de · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.655 before conversion and 0.716 after — it rose by 0.061. Neighbour-to-neighbour the worst pair went 0.709 → 0.772. (The earlier render, with segment 1 left raw, scores 0.734 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved -0.251 in the original and -0.064 after conversion — 25 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Pride, -0.455 became -0.764.

Quality. Mean predicted overall quality across the segments went 2.74 → 2.81 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.655 → 0.716 +0.061identity cos neighbours 0.709 → 0.772d_b rescored -0.251 → -0.064d_a rescored -0.455 → -0.764d_a mined -0.455d_b mined -0.253min_cos_consec (site) 0.7294min_cos_anchor (site) 0.5979dataset emolialang despeaker DE_B00002_S07600total 20.1schain gain +2.0 dBseam step 1.3 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, slightly relaxed, fairly steady
(pride, triumph · normal-paced, energised, no disfluency, formal) Jetzt lernen wir die Buchstaben vom Alphabet.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as pride, triumph; style: formal, narration; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 3.2s, DE.
DE_B00002_S07600_W000058 · in -23.1 dBFS · gain +3.1 dB · emolia-00056
(teasing · measured, normally alert, almost no disfluency, storytelling) Mit Buchstaben kannst du Wörter schreiben. Jetzt lernen wir die Buchstaben K bis O. Lass uns schnell anfangen.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as teasing; style: storytelling, narration; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.5/10; 7.3s, DE.
DE_B00002_S07600_W000059 · in -23.2 dBFS · gain +3.2 dB · emolia-00056
(pride, triumph, teasing · normal-paced, normally alert, no disfluency, formal) Der erste Buchstabe, den wir lernen, ist der Buchstabe K.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pride, triumph, teasing; style: formal, storytelling; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.0/10; 3.4s, DE.
DE_B00002_S07600_W000060 · in -23.6 dBFS · gain +3.6 dB · emolia-00056
(measured, normally alert, no disfluency, storytelling) Der erste Buchstabe vom Wort Käse ist ein K.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; no dominant emotion; style: storytelling, narration; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.0/10; 3.5s, DE.
DE_B00002_S07600_W000061 · in -22.9 dBFS · gain +2.9 dB · emolia-00056
(normal-paced, normally alert, no disfluency, formal) Der nächste Buchstabe, den wir lernen, ist der Buchstabe L.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.2/10; 3.6s, DE.
DE_B00002_S07600_W000062 · in -21.1 dBFS · gain +1.1 dB · emolia-00056
Emotional Numbness ↓  /  Concentrationidentity +0.72 emotion 101 %   c-emolia-AB2 · #18

This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Concentration up — by at least 0.25 each.

The chain starts with Concentration below average — 0.35, lower than 65 % of clips in this corpus — and ends with it clearly present at 0.68, higher than 68 % of clips in this corpus. That is a total rise of 0.33.

At the same time Emotional Numbness goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.13, then +0.08, then -0.06, then +0.18 — not a clean run: step 3 moves back the other way by 0.06 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.12 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.07 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.12, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 24 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.119 before conversion and 0.600 after — it rose by 0.719. Neighbour-to-neighbour the worst pair went -0.114 → 0.442. (The earlier render, with segment 1 left raw, scores 0.426 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.334 in the original and +0.339 after conversion — 101 % of the delta retained, which is essentially all of it. On the other named axis, Emotional Numbness, -0.281 became -0.233.

Quality. Mean predicted overall quality across the segments went 2.49 → 2.67 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 -0.119 → 0.600 +0.719identity cos neighbours -0.114 → 0.442d_b rescored +0.334 → +0.339d_a rescored -0.281 → -0.233d_a mined -0.281d_b mined 0.334min_cos_consec (site) -0.0744min_cos_anchor (site) -0.1174dataset emolialang enspeaker EN_CMxzIUnXLDAtotal 22.5schain gain +2.6 dBseam step 1.4 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady
(some disfluency, somewhat unclear, casual, monologue) And, (ahem) uh, there was, uh, they have been in negative press, of course, but.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 1.3/10; 4.8s, EN.
EN_CMxzIUnXLDA_W000029 · in -18.8 dBFS · gain -1.2 dB · emolia-01913
(relief, emotional numbness · little disfluency, clear, formal, authoritative) With the approach that they're taking, I'm not so much concerned.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as relief, emotional numbness; style: formal, authoritative; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.7/10; 3.2s, EN.
EN_CMxzIUnXLDA_W000030 · in -19.6 dBFS · gain -0.4 dB · emolia-01913
(emotional numbness · some disfluency, average clarity, monologue, casual) We'll have one vote, equal vote, and all the decisions that are going to be made on a governance, from a governance perspective.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: monologue, casual; average recording, no background noise; genuineness 3.1/6; vocal-burst blend 1.5/10; 6.5s, EN.
EN_CMxzIUnXLDA_W000033 · in -19.7 dBFS · gain -0.3 dB · emolia-01913
(pride, triumph · some disfluency, average clarity, casual, playful) (low mumble) Uh, are going to be made democratically. So right now it's like a federated network.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as pride, triumph; style: casual, playful; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 2.8/10; 3.9s, EN.
EN_CMxzIUnXLDA_W000034 · in -18.5 dBFS · gain -1.5 dB · emolia-01913
(little disfluency, average clarity, monologue, formal) (low mumble) Uhm, with all of the association members. And over time you can move towards a permissionless system.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 1.1/10; 4.8s, EN.
EN_CMxzIUnXLDA_W000035 · in -18.5 dBFS · gain -1.5 dB · emolia-01913
Pain ↓  /  Doubtidentity −0.00 emotion 29 %   c-emolia-AB2 · #19

This chain comes from the two-sided rule: it only counts if both emotions move — Pain down and Doubt up — by at least 0.25 each.

The chain starts with Doubt around average — 0.55, higher than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.37.

At the same time Pain goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.20 (lower than 80 % of clips in this corpus), a change of -0.52. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.11, then +0.03, then +0.23 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 38 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.721 before conversion and 0.716 after — it fell by 0.005. Neighbour-to-neighbour the worst pair went 0.721 → 0.716. (The earlier render, with segment 1 left raw, scores 0.695 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.368 in the original and +0.105 after conversion — 29 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Pain, -0.520 became +0.369.

Quality. Mean predicted overall quality across the segments went 3.02 → 3.20 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.721 → 0.716 -0.005identity cos neighbours 0.721 → 0.716d_b rescored +0.368 → +0.105d_a rescored -0.520 → +0.369d_a mined -0.520d_b mined 0.368min_cos_consec (site) 0.8081min_cos_anchor (site) 0.8052dataset emolialang zhspeaker ZH_B00017_S09107total 37.1schain gain +0.5 dBseam step 1.6 dBcrossfades 150/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(measured, no disfluency, narration, monologue) 不同治理主体之间呢还要合理制约,避免任何一方过度膨胀。
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 2.5/10; 5.6s, ZH.
ZH_B00017_S09107_W000024 · in -20.9 dBFS · gain +0.9 dB · emolia-03448
(normal-paced, some disfluency, monologue, narration) 比如说校务会虽然是重大决策机构,但每年的教师代表大会要对校务委员们进行满意度投票,还要对校长进行信任度投票。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 4.7/10; 11.4s, ZH.
ZH_B00017_S09107_W000025 · in -20.4 dBFS · gain +0.4 dB · emolia-03448
(longing · measured, no disfluency, narration, monologue) 这就是治理主体之间的一种微妙的制约,使得校长的日常行为不能够太随意。再比如,学术委员会在十一学校全部由教师代表组成,校务会成员不能进入。
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing; style: narration, monologue; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 3.6/10; 14.1s, ZH.
ZH_B00017_S09107_W000026 · in -21.0 dBFS · gain +1.0 dB · emolia-03448
(doubt, astonishment surprise · measured, no disfluency, monologue, narration) 这些教师学术水平高超,道德境界得到公认,同时又不在管理的序列里。
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, astonishment surprise; style: monologue, narration; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 3.3/10; 6.6s, ZH.
ZH_B00017_S09107_W000027 · in -20.2 dBFS · gain +0.2 dB · emolia-03448
Astonishment Surprise ↓  /  Interestidentity +0.02 emotion 99 %   c-emolia-AB2 · #20

This chain comes from the two-sided rule: it only counts if both emotions move — Astonishment Surprise down and Interest up — by at least 0.25 each.

The chain starts with Interest below average — 0.37, lower than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.63.

At the same time Astonishment Surprise goes the other way, from 0.66 (higher than 66 % of clips in this corpus) to 0.23 (lower than 78 % of clips in this corpus), a change of -0.43. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.22, then +0.19, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.98 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.98 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.98), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 57 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.784 before conversion and 0.801 after — it rose by 0.016. Neighbour-to-neighbour the worst pair went 0.897 → 0.917. (The earlier render, with segment 1 left raw, scores 0.633 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.628 in the original and +0.620 after conversion — 99 % of the delta retained, which is essentially all of it. On the other named axis, Astonishment Surprise, -0.434 became -0.299.

Quality. Mean predicted overall quality across the segments went 3.09 → 3.25 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.784 → 0.801 +0.016identity cos neighbours 0.897 → 0.917d_b rescored +0.628 → +0.620d_a rescored -0.434 → -0.299d_a mined -0.434d_b mined 0.628min_cos_consec (site) 0.9814min_cos_anchor (site) 0.9753dataset emolialang enspeaker EN_9ko2VaebzYctotal 55.7schain gain +1.6 dBseam step 1.0 dBcrossfades 150/100/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fairly steady, light breath, formal, newsreading) According to Germano, the Dzogchen tradition first appeared in the first half of the 9th century, with a series of short texts attributed to Indian saints
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 8.1s, EN.
EN_9ko2VaebzYc_W000050 · in -15.6 dBFS · gain -4.4 dB · emolia-00454
(steady, light breath, newsreading, formal) The mind series reflect the teachings of early Dzogchen, which rejected all forms of practice, and asserted that striving for liberation would simply create more delusion
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 8.8s, EN.
EN_9ko2VaebzYc_W000051 · in -15.1 dBFS · gain -4.9 dB · emolia-00454
(concentration, contemplation · steady, minimal breath, newsreading, authoritative) One has simply to recognize the nature of one's own mind, which is naturally empty, stong pa, luminous, odd gsal ba, and pure. According to Germano, its characteristic language, which is marked by naturalism and negation, is already pronounced in some Indian tantras.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, contemplation; style: newsreading, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 16.2s, EN.
EN_9ko2VaebzYc_W000052 · in -15.5 dBFS · gain -4.5 dB · emolia-00454
(interest, concentration, awe · steady, minimal breath, newsreading, formal) Nevertheless, these texts are still inextricably bound up with Tantric Mahayoga, with its visualizations of deities and mandals, and complex initiations.During the 9th and 10th centuries these texts, which represent the dominant form of the tradition in the 9th and 10th centuries, were gradually transformed into full-fledged tantras, culminating in the Kulairaja Tantra' Khun Bhide Arjeel Poh, "'the all-creating king'
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as interest, concentration, awe; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 23.1s, EN.
EN_9ko2VaebzYc_W000053 · in -15.4 dBFS · gain -4.6 dB · emolia-00454