c-emolia-B1 — voice-corrected

Corpus emolia in isolation, rule B1.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_c-emolia-B1.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
62segments re-voiced
0.669 → 0.731median worst-to-anchor identity cosine
78 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Contemplation(unconstrained axis: Triumph)identity +0.46 emotion 60 %   c-emolia-B1 · #1

This chain comes from the one-sided rule: only Contemplation had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Contemplation clearly present — 0.58, higher than 58 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.37.

Nothing was asked of the other axis, and in fact Triumph drifts down from 0.82 to 0.51 (-0.31), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.11, then -0.01, then +0.06, then +0.21 — not a clean run: step 2 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 34 s · de · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.310 before conversion and 0.769 after — it rose by 0.460. Neighbour-to-neighbour the worst pair went 0.361 → 0.803. (The earlier render, with segment 1 left raw, scores 0.652 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.365 in the original and +0.220 after conversion — 60 % of the delta retained. On the other named axis, Triumph, -0.312 became -0.222.

Quality. Mean predicted overall quality across the segments went 2.83 → 2.92 (+0.09) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.310 → 0.769 +0.460identity cos neighbours 0.361 → 0.803d_b rescored +0.365 → +0.220d_a rescored -0.312 → -0.222d_a mined -0.312d_b mined 0.365min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_YwdBKmM2CpEtotal 32.4schain gain +1.4 dBseam step 3.1 dBcrossfades 100/100/100/150 ms
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(normal-paced, some disfluency, average clarity, casual) Wir haben ja damals auf die Games im Unterricht-Seite deswegen auch extra die pädagogische Empfehlung ab 14 geschrieben, obwohl das Spiel ja eigentlich ab 12 beigegeben ist. Und dann noch mal dazu geschrieben, wieso, (ahem) genau aus diesem Grund, ja.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 4.0/6; vocal-burst blend 4.5/10; 11.3s, DE.
DE_YwdBKmM2CpE_W000083 · in -18.7 dBFS · gain -1.3 dB · emolia-00003
(normal-paced, frequent disfluency, somewhat unclear, conversational) Genau, also das ist auch, äh, (low mumble) interessant, äh, (low mumble) mit den pädagogischen Empfehlungen. Also wir haben in den Spielen ja USK,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: conversational, monologue; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 0.0/10; 7.0s, DE.
DE_YwdBKmM2CpE_W000084 · in -18.9 dBFS · gain -1.1 dB · emolia-00003
(distress · measured, frequent disfluency, somewhat unclear, monologue) dafür dienen, dass man es Kinder jetzt nicht verstören soll und, und das, (low mumble) ähm, da auch aus.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as distress; style: monologue, casual; average recording, no background noise; genuineness 3.5/6; vocal-burst blend 0.0/10; 4.3s, DE.
DE_YwdBKmM2CpE_W000085 · in -16.0 dBFS · gain -4.0 dB · emolia-00003
(normal-paced, frequent disfluency, somewhat unclear) Sag ich mal, Jugendschutzperspektive, 呃, (low mumble) grad das Thema Gewalt, 呃, (low mumble)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; average recording, no background noise; genuineness 3.5/6; vocal-burst blend 0.0/10; 4.1s, DE.
DE_YwdBKmM2CpE_W000086 · in -17.1 dBFS · gain -2.9 dB · emolia-00003
(contemplation, doubt · normal-paced, some disfluency, somewhat unclear, monologue) in welcher Form das irgendwie, (low mumble) ja, schwierig sein könnte. Aber es ist keine pädagogische Empfehlung und, (low mumble) ähm.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as contemplation, doubt; style: monologue, casual; good recording, quiet background; genuineness 4.0/6; vocal-burst blend 1.1/10; 6.4s, DE.
DE_YwdBKmM2CpE_W000087 · in -20.4 dBFS · gain +0.4 dB · emolia-00003
Contemplation(unconstrained axis: Fatigue Exhaustion)identity −0.01 emotion 64 %   c-emolia-B1 · #2

This chain comes from the one-sided rule: only Contemplation had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Contemplation clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.31.

Nothing was asked of the other axis, and in fact Fatigue Exhaustion drifts down from 0.88 to 0.45 (-0.43), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.07, then +0.18, then +0.06 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 49 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.937 before conversion and 0.931 after — it fell by 0.006. Neighbour-to-neighbour the worst pair went 0.949 → 0.933. (The earlier render, with segment 1 left raw, scores 0.887 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.313 in the original and +0.199 after conversion — 64 % of the delta retained. On the other named axis, Fatigue Exhaustion, -0.433 became -0.445.

Quality. Mean predicted overall quality across the segments went 3.19 → 3.40 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.937 → 0.931 -0.006identity cos neighbours 0.949 → 0.933d_b rescored +0.313 → +0.199d_a rescored -0.433 → -0.445d_a mined -0.433d_b mined 0.313min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00000_S07855total 48.4schain gain +1.9 dBseam step 0.6 dBcrossfades 150/100/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady
(normal-paced, narration, monologue) 他们生产的只不过是在松软石粒和沙子土壤上到处都能生产的优良普通葡萄酒。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 3.7/10; 8.4s, ZH.
ZH_B00000_S07855_W000114 · in -20.0 dBFS · gain +0.0 dB · emolia-03273
(measured, monologue, formal) 除了强度和卫生外,其他无阻称道,一国的普通土地只是同这样的葡萄园才能进行竞争。同具有特殊品质的葡萄园,它显然无法竞争。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 0.2/6; vocal-burst blend 3.0/10; 12.1s, ZH.
ZH_B00000_S07855_W000115 · in -21.2 dBFS · gain +1.2 dB · emolia-03273
(measured, monologue, narration) 例如,生产特殊味道的葡萄酒的土地,葡萄比其他任何果树更易受土壤性质不同的影响。一般认为,葡萄从某种土壤获得一种滋味,任何培育或管理办法都不能做到。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 4.5/10; 15.2s, ZH.
ZH_B00000_S07855_W000116 · in -20.0 dBFS · gain -0.0 dB · emolia-03273
(contemplation, pride · measured, monologue, narration) 这种味道不论是真实的,还是想象的是少数葡萄园的产品有时所特有的,有时它扩大到一个小地区的大部分地方,有时扩大到一个省的绝大部分地区。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, pride; style: monologue, narration; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 4.8/10; 13.1s, ZH.
ZH_B00000_S07855_W000117 · in -18.7 dBFS · gain -1.3 dB · emolia-03273
Pain(unconstrained axis: Emotional Numbness)identity −0.07 emotion 48 %   c-emolia-B1 · #3

This chain comes from the one-sided rule: only Pain had to get where it was going, by at least 0.50. The other emotion was left completely free.

The chain starts with Pain below average — 0.35, lower than 65 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.52.

Nothing was asked of the other axis, and in fact Emotional Numbness drifts down from 0.91 to 0.79 (-0.12), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.17, then +0.16 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 43 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.784 before conversion and 0.711 after — it fell by 0.073. Neighbour-to-neighbour the worst pair went 0.784 → 0.711. (The earlier render, with segment 1 left raw, scores 0.588 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.524 in the original and +0.253 after conversion — 48 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Emotional Numbness, -0.115 became -0.001.

Quality. Mean predicted overall quality across the segments went 2.97 → 3.11 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.784 → 0.711 -0.073identity cos neighbours 0.784 → 0.711d_b rescored +0.524 → +0.253d_a rescored -0.115 → -0.001d_a mined -0.115d_b mined 0.524min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00064_S04083total 41.8schain gain +1.2 dBseam step 0.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(emotional numbness · fairly steady, no disfluency, formal, authoritative) For narrative types by definition have consistent structure, and follow an existing
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.2/10; 5.6s, EN.
EN_B00064_S04083_W000068 · in -14.7 dBFS · gain -5.3 dB · emolia-01469
(contentment · steady, no disfluency, newsreading, formal) For childhood is a social group where children teach, learn and share their own traditions, flourishing
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contentment; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 16.4s, EN.
EN_B00064_S04083_W000069 · in -14.7 dBFS · gain -5.3 dB · emolia-01469
(emotional numbness, concentration, malevolence malice · steady, almost no disfluency, newsreading, formal) With an increasingly theoretical sophistication of the social sciences, it has become evident
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, concentration, malevolence malice; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.1s, EN.
EN_B00064_S04083_W000070 · in -14.7 dBFS · gain -5.3 dB · emolia-01469
(fairly steady, no disfluency, formal, newsreading) The Fairy Tale Snow White is now offered in multiple media forms for both children and adults, including a television show, a video game, and a programming language
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 9.3s, EN.
EN_B00064_S04083_W000071 · in -14.0 dBFS · gain -6.0 dB · emolia-01469
Disgust(unconstrained axis: Fear)identity −0.16 emotion 97 %   c-emolia-B1 · #4

This chain comes from the one-sided rule: only Disgust had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Disgust around average — 0.53, higher than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.43.

Nothing was asked of the other axis, and in fact Fear drifts down from 0.81 to 0.27 (-0.53), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.12, then +0.09, then +0.21 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 38 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.952 before conversion and 0.790 after — it fell by 0.162. Neighbour-to-neighbour the worst pair went 0.948 → 0.863. (The earlier render, with segment 1 left raw, scores 0.678 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disgust moved +0.427 in the original and +0.415 after conversion — 97 % of the delta retained, which is essentially all of it. On the other named axis, Fear, -0.533 became -0.631.

Quality. Mean predicted overall quality across the segments went 2.99 → 3.14 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.952 → 0.790 -0.162identity cos neighbours 0.948 → 0.863d_b rescored +0.427 → +0.415d_a rescored -0.533 → -0.631d_a mined -0.533d_b mined 0.427min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_MlD9ZsWreJEtotal 36.8schain gain +1.5 dBseam step 0.4 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-bright, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, steady, clear
(normal-paced, no disfluency, moderate pitch range, formal) The architecture was greeted by Seattle residents with a mixture of acclaim for Gerry and derision for this particular edifice,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 7.8s, EN.
EN_MlD9ZsWreJE_W000078 · in -14.7 dBFS · gain -5.3 dB · emolia-02233
(awe, pride · normal-paced, almost no disfluency, moderate pitch range, newsreading) Remarked British-born, Seattle-based writer Jonathan Rabin, "...has created some wonderful buildings, like the Guggenheim Museum in Bilbao, but his Seattle effort, the Experience Music Project, is not one of them."
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, pride; style: newsreading, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 13.3s, EN.
EN_MlD9ZsWreJE_W000079 · in -13.1 dBFS · gain -6.9 dB · emolia-02233
(emotional numbness · normal-paced, no disfluency, moderate pitch range, formal) New York Times architecture critic Herbert Mushamp described it as, "...something that crawled out of the sea, rolled over, and died
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.8s, EN.
EN_MlD9ZsWreJE_W000080 · in -17.8 dBFS · gain -2.2 dB · emolia-02233
(disgust · measured, no disfluency, fairly narrow pitch, formal) Forbes magazine called it one of the world's ten ugliest buildings. Others describe it as a
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 8.6s, EN.
EN_MlD9ZsWreJE_W000081 · in -17.5 dBFS · gain -2.5 dB · emolia-02233
Pride(unconstrained axis: Doubt)identity +0.31 emotion 130 %   c-emolia-B1 · #5

This chain comes from the one-sided rule: only Pride had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Pride clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.26.

Nothing was asked of the other axis, and in fact Doubt drifts down from 0.97 to 0.55 (-0.42), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are -0.18, then +0.22, then +0.05, then +0.18 — not a clean run: step 1 moves back the other way by 0.18 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.36 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.36 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 50 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.291 before conversion and 0.605 after — it rose by 0.315. Neighbour-to-neighbour the worst pair went 0.205 → 0.617. (The earlier render, with segment 1 left raw, scores 0.357 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.265 in the original and +0.345 after conversion — 130 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Doubt, -0.417 became -0.472.

Quality. Mean predicted overall quality across the segments went 2.65 → 3.00 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.291 → 0.605 +0.315identity cos neighbours 0.205 → 0.617d_b rescored +0.265 → +0.345d_a rescored -0.417 → -0.472d_a mined -0.416d_b mined 0.265min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_EdahOwgKcOktotal 48.1schain gain +4.5 dBseam step 1.9 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, normal-paced, normally alert, slightly relaxed
(doubt, contemplation, fear · average clarity, monologue, casual) And they may not think about how the customer sees that service from the business point of view. They'll think I'm just going to expose out this technical capability with an API. And of course they're going to want to use it because we use it.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as doubt, contemplation, fear; style: monologue, casual; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 3.3/10; 11.1s, EN.
EN_EdahOwgKcOk_W000042 · in -21.0 dBFS · gain +1.0 dB · emolia-00667
(interest, concentration · somewhat unclear, whispered, monologue) It's a different ball game. Once you cross that boundary between internal systems and external systems, start making things customer facing. You really got to think about the (low mumble) business service you're providing with that API.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, concentration; style: whispered, monologue; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 3.4/10; 11.6s, EN.
EN_EdahOwgKcOk_W000043 · in -24.1 dBFS · gain +4.0 dB · emolia-00667
(average clarity, casual, conversational) Uh, (low mumble) and happen to expose it to others and as we've gone along,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 1.3/10; 4.3s, EN.
EN_EdahOwgKcOk_W000045 · in -20.6 dBFS · gain +0.6 dB · emolia-00667
(contemplation, concentration, hope enthusiasm optimism · average clarity, casual, monologue) You know, we feel like the growth of that API kind of plateaued because it was so tightly tied to how we saw the product and needed to start thinking to how our customers wanted to interact with our API products and make changes to address that gap.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as contemplation, concentration, hope enthusiasm optimism; style: casual, monologue; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 5.8/10; 16.0s, EN.
EN_EdahOwgKcOk_W000046 · in -22.9 dBFS · gain +2.9 dB · emolia-00667
(pride, relief, thankfulness gratitude · average clarity, casual, monologue) Yeah, (ahem) a lot of people go through that journey. We had the same challenge when I was at City with our APIs.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as pride, relief, thankfulness gratitude; style: casual, monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 2.9/10; 5.9s, EN.
EN_EdahOwgKcOk_W000047 · in -21.9 dBFS · gain +1.9 dB · emolia-00667
Relief(unconstrained axis: Confusion)identity +0.05 emotion 116 %   c-emolia-B1 · #6

This chain comes from the one-sided rule: only Relief had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Relief clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.34.

Nothing was asked of the other axis, and in fact Confusion drifts down from 0.94 to 0.75 (-0.19), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.09 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 53 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.785 before conversion and 0.830 after — it rose by 0.045. Neighbour-to-neighbour the worst pair went 0.785 → 0.830. (The earlier render, with segment 1 left raw, scores 0.734 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.338 in the original and +0.392 after conversion — 116 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Confusion, -0.188 became -0.133.

Quality. Mean predicted overall quality across the segments went 2.86 → 3.18 (+0.32) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.785 → 0.830 +0.045identity cos neighbours 0.785 → 0.830d_b rescored +0.338 → +0.392d_a rescored -0.188 → -0.133d_a mined -0.188d_b mined 0.338min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00003_S00684total 52.5schain gain +2.4 dBseam step 1.7 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, balanced body, average recording, measured, slightly relaxed, fairly steady, moderate pitch range, light breath
(confusion, emotional numbness · normally alert, some disfluency, slurred, monologue) 最好的两家最顶级的,我查一下网红第一名酒店,爱迪逊跟那个保利的对吧?最好的酒店。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, emotional numbness; style: monologue, didactic; average recording, no background noise; genuineness 2.7/6; vocal-burst blend 2.1/10; 8.4s, ZH.
ZH_B00003_S00684_W000114 · in -25.5 dBFS · gain +5.5 dB · emolia-03306
(jealousy and envy, contempt, contemplation · subdued, some disfluency, somewhat unclear, monologue) (ahem) 也就五星级酒店它有分啊,有有一千块钱的,有两千块钱的,也有三千块钱的,我们就只三千块钱。所以安利的这个旅游我我感觉是几个点。第一个是高端,很高端,很奢华,同时很精致,很注重细节。你看我们这个是这个酒店的那个湖边泳池啊,然后我们到了酒店之后,所有人这个把这个礼物礼品。
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as jealousy and envy, contempt, contemplation; style: monologue, didactic; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 6.8/10; 30.0s, ZH.
ZH_B00003_S00684_W000115 · in -22.6 dBFS · gain +2.6 dB · emolia-03306
(relief, sourness, contentment · normally alert, frequent disfluency, somewhat unclear, monologue) (ahem) (ahem) 我们算了一下,包括放在放在现场,你可以拿的,还有送到你家里的,已经寄到你家里的,就这个呃这个这个啊早餐机,他送到你家里,我都算了一下,包括还有很多产品,包括很多。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, sourness, contentment; style: monologue, casual; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 6.4/10; 14.5s, ZH.
ZH_B00003_S00684_W000116 · in -22.9 dBFS · gain +2.9 dB · emolia-03306
Impatience and Irritability(unconstrained axis: Disappointment)identity +0.18 emotion 92 %   c-emolia-B1 · #7

This chain comes from the one-sided rule: only Impatience and Irritability had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Impatience and Irritability clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.29.

Nothing was asked of the other axis, and in fact Disappointment drifts down from 0.99 to 0.91 (-0.08), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.17, then +0.09, then +0.02 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 32 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.412 before conversion and 0.593 after — it rose by 0.181. Neighbour-to-neighbour the worst pair went 0.510 → 0.774. (The earlier render, with segment 1 left raw, scores 0.529 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.286 in the original and +0.262 after conversion — 92 % of the delta retained, which is essentially all of it. On the other named axis, Disappointment, -0.082 became -0.069.

Quality. Mean predicted overall quality across the segments went 2.93 → 3.03 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.412 → 0.593 +0.181identity cos neighbours 0.510 → 0.774d_b rescored +0.286 → +0.262d_a rescored -0.082 → -0.069d_a mined -0.082d_b mined 0.286min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00067_S00928total 30.9schain gain +1.8 dBseam step 1.7 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult feminine voice · no background noise, energised
(disappointment, pain, doubt · measured, slightly relaxed, moderately variable, dramatic) We just thought it would be a good solution to the situation. Nothing more than that.
full caption & clip details
An adult feminine voice; delivery is energised, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as disappointment, pain, doubt; style: dramatic, storytelling; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 1.6/10; 5.8s, ZH.
ZH_B00067_S00928_W000601 · in -19.8 dBFS · gain -0.2 dB · emolia-03950
(contempt, disgust, emotional numbness · brisk, slightly tense, moderately variable, cartoonish) About a property investment. It's not about profit and loss. It's about a human being.
full caption & clip details
An adult feminine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; very clear, no disfluency, very wide pitch range, light breath; affect is positive, slightly dominant, guarded; reads as contempt, disgust, emotional numbness; style: cartoonish, ranting; average recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.7/10; 6.2s, ZH.
ZH_B00067_S00928_W000602 · in -18.0 dBFS · gain -2.0 dB · emolia-03950
(disappointment, distress, disgust · brisk, slightly tense, volatile, casual) I think you've lost all your human feelings. You're like a machine program to make money. Can't you see how discussing it is? (contented sigh) What's wrong with you?
full caption & clip details
An adult feminine voice; delivery is energised, brisk, slightly tense, volatile; timbre is slightly cool, slightly bright, very rough, full; clear, some disfluency, very wide pitch range, normal breath; affect is negative, slightly dominant, guarded; reads as disappointment, distress, disgust; style: casual, storytelling; average recording, no background noise; mildly explicit content; genuineness 1.7/6; vocal-burst blend 1.2/10; 11.6s, ZH.
ZH_B00067_S00928_W000603 · in -19.8 dBFS · gain -0.2 dB · emolia-03950
(impatience and irritability, anger, teasing · brisk, slightly tense, moderately variable, storytelling) If you can afford to pay her one thousand five hundred pounds a month, then just pay it fdon't 't expect to make a profit out of it.
full caption & clip details
A child feminine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, no disfluency, very wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as impatience and irritability, anger, teasing; style: storytelling, cartoonish; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.1/10; 8.0s, ZH.
ZH_B00067_S00928_W000604 · in -17.7 dBFS · gain -2.3 dB · emolia-03950
Emotional Numbness(unconstrained axis: Hope Enthusiasm Optimism)identity +0.00 emotion 56 %   c-emolia-B1 · #8

This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Emotional Numbness around average — 0.55, higher than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.37.

Nothing was asked of the other axis, and in fact Hope Enthusiasm Optimism drifts down from 0.93 to 0.38 (-0.55), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.18, then +0.20, then -0.01 — not a clean run: step 3 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 27 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.655 before conversion and 0.656 after — it rose by 0.002. Neighbour-to-neighbour the worst pair went 0.793 → 0.631. (The earlier render, with segment 1 left raw, scores 0.575 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.370 in the original and +0.206 after conversion — 56 % of the delta retained. On the other named axis, Hope Enthusiasm Optimism, -0.549 became -0.470.

Quality. Mean predicted overall quality across the segments went 2.69 → 2.87 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.655 → 0.656 +0.002identity cos neighbours 0.793 → 0.631d_b rescored +0.370 → +0.206d_a rescored -0.549 → -0.470d_a mined -0.549d_b mined 0.370min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_l1wwbQNMbj8total 25.8schain gain +2.8 dBseam step 1.4 dBcrossfades 100/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, balanced body, slightly relaxed
(hope enthusiasm optimism, affection · measured, subdued, steady, whispered) If life on the water appeals to you, hey, come on and check out the gentry. So let's continue on around the gentry and you'll see entering the picture right there.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as hope enthusiasm optimism, affection; style: whispered, monologue; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.8/10; 11.5s, EN.
EN_l1wwbQNMbj8_W000018 · in -22.1 dBFS · gain +2.1 dB · emolia-02331
(normal-paced, normally alert, fairly steady, casual) Are the community boat ramps, those are accessible to the entire Hansel community and.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 3.0/6; vocal-burst blend 1.8/10; 5.9s, EN.
EN_l1wwbQNMbj8_W000019 · in -20.9 dBFS · gain +0.9 dB · emolia-02331
(emotional numbness · normal-paced, normally alert, fairly steady, monologue) Through the gentry itself, so easy for residents to get to.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: monologue, casual; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 2.0/10; 3.5s, EN.
EN_l1wwbQNMbj8_W000020 · in -18.9 dBFS · gain -1.1 dB · emolia-02331
(emotional numbness · normal-paced, normally alert, fairly steady, monologue) The gentry is a relatively small community and you can take in its entirety through this video.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: monologue, formal; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.0/10; 5.3s, EN.
EN_l1wwbQNMbj8_W000021 · in -17.3 dBFS · gain -2.7 dB · emolia-02331
Fatigue Exhaustion(unconstrained axis: Intoxication Altered States of Consciousness)identity −0.04 emotion 189 %   c-emolia-B1 · #9

This chain comes from the one-sided rule: only Fatigue Exhaustion had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Fatigue Exhaustion around average — 0.57, higher than 57 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.27.

Nothing was asked of the other axis, and in fact Intoxication Altered States of Consciousness drifts down from 0.87 to 0.28 (-0.60), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.10, then -0.25, then +0.25, then +0.17 — not a clean run: step 2 moves back the other way by 0.25 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.82 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.82 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 23 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.788 before conversion and 0.750 after — it fell by 0.038. Neighbour-to-neighbour the worst pair went 0.754 → 0.750. (The earlier render, with segment 1 left raw, scores 0.709 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.268 in the original and +0.507 after conversion — 189 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Intoxication Altered States of Consciousness, -0.596 became -0.629.

Quality. Mean predicted overall quality across the segments went 2.72 → 2.90 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.788 → 0.750 -0.038identity cos neighbours 0.754 → 0.750d_b rescored +0.268 → +0.507d_a rescored -0.596 → -0.629d_a mined -0.596d_b mined 0.268min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00029_S01551total 21.8schain gain +2.2 dBseam step 1.7 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(measured, some disfluency, average clarity, casual) 当然你在看的时候就会比较的模糊,ok.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, playful; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 1.2/10; 4.1s, ZH.
ZH_B00029_S01551_W001118 · in -18.1 dBFS · gain -1.9 dB · emolia-03563
(measured, frequent disfluency, slurred, conversational) (low mumble) 那么下面就有一个人在吐槽说。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; no dominant emotion; style: conversational, casual; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 1.2/10; 3.6s, ZH.
ZH_B00029_S01551_W001119 · in -19.6 dBFS · gain -0.4 dB · emolia-03563
(fast, no disfluency, clear, authoritative) 搜寻前几平的文章,他可以帮助你理解,但是他讲的不一定是对的。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; average recording, no background noise; genuineness 2.0/6; vocal-burst blend 3.2/10; 4.4s, ZH.
ZH_B00029_S01551_W001120 · in -19.0 dBFS · gain -1.0 dB · emolia-03563
(measured, some disfluency, average clarity, conversational) 所以我们要去看文件,我们看文件,那我们来看four of的文件。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational, monologue; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 2.6/10; 4.9s, ZH.
ZH_B00029_S01551_W001121 · in -18.0 dBFS · gain -2.0 dB · emolia-03563
(measured, some disfluency, average clarity, authoritative) 那我是已经切换到英文版,我先切换回简体中文,让大家稍微看一下。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, monologue; average recording, no background noise; genuineness 2.6/6; vocal-burst blend 3.6/10; 5.6s, ZH.
ZH_B00029_S01551_W001122 · in -18.6 dBFS · gain -1.4 dB · emolia-03563
Embarrassment(unconstrained axis: Hope Enthusiasm Optimism)identity −0.01 emotion 102 %   c-emolia-B1 · #10

This chain comes from the one-sided rule: only Embarrassment had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Embarrassment clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.25.

Nothing was asked of the other axis, and in fact Hope Enthusiasm Optimism drifts down from 0.97 to 0.63 (-0.34), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 29 s · zh · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.937 before conversion and 0.929 after — it fell by 0.007. Neighbour-to-neighbour the worst pair went 0.937 → 0.929. (The earlier render, with segment 1 left raw, scores 0.852 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.249 in the original and +0.253 after conversion — 102 % of the delta retained, which is essentially all of it. On the other named axis, Hope Enthusiasm Optimism, -0.343 became -0.330.

Quality. Mean predicted overall quality across the segments went 3.00 → 3.19 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.937 → 0.929 -0.007identity cos neighbours 0.937 → 0.929d_b rescored +0.249 → +0.253d_a rescored -0.343 → -0.330d_a mined -0.344d_b mined 0.249min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00050_S00428total 28.3schain gain +3.2 dBseam step 0.3 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a child feminine voice · slightly cool, slightly bright, fairly smooth, thin, average recording, no background noise, energised, slightly relaxed
(hope enthusiasm optimism, pleasure ecstasy, elation · brisk, cartoonish, dramatic) 这两个和这两个呢是一样的,可以相互替换。那么唯一的区别呢是这个handtest和这个色比起这个前面的呢,是有一点这个口语的语气,对吧?那么这个huntaget是什么意思呢?是他们两个人的尊称,对吧?
full caption & clip details
A child feminine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, pleasure ecstasy, elation; style: cartoonish, dramatic; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 6.0/10; 15.8s, ZH.
ZH_B00050_S00428_W000003 · in -21.6 dBFS · gain +1.6 dB · emolia-03780
(embarrassment, sexual lust, disgust · fast, cartoonish, dramatic) 比如说爷爷,那么爷爷的后面是加这个gay啊,那么是什么意思呢?这个a gay hunter和ay呢,是给的意思,对吧?给给妈妈写信给爸爸打电话的,给。那么这个。
full caption & clip details
A child feminine voice; delivery is energised, fast, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as embarrassment, sexual lust, disgust; style: cartoonish, dramatic; average recording, no background noise; genuineness 2.8/6; vocal-burst blend 5.7/10; 12.6s, ZH.
ZH_B00050_S00428_W000004 · in -20.9 dBFS · gain +0.9 dB · emolia-03780
Hope Enthusiasm Optimism(unconstrained axis: Sexual Lust)identity −0.07 emotion 29 %   c-emolia-B1 · #11

This chain comes from the one-sided rule: only Hope Enthusiasm Optimism had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Hope Enthusiasm Optimism clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.26.

Nothing was asked of the other axis, and in fact Sexual Lust barely moves at all, sitting near 0.97 throughout.

It takes 5 clips to get there. Clip to clip the moves are +0.06, then +0.02, then +0.04, then +0.14 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.77 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.77 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 26 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.523 before conversion and 0.451 after — it fell by 0.072. Neighbour-to-neighbour the worst pair went 0.523 → 0.451. (The earlier render, with segment 1 left raw, scores 0.417 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.255 in the original and +0.073 after conversion — 29 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Sexual Lust, -0.043 became -0.048.

Quality. Mean predicted overall quality across the segments went 2.56 → 2.76 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.523 → 0.451 -0.072identity cos neighbours 0.523 → 0.451d_b rescored +0.255 → +0.073d_a rescored -0.043 → -0.048d_a mined -0.043d_b mined 0.256min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_hj-Jo9BzTa8total 25.1schain gain +2.3 dBseam step 1.9 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a child feminine voice · neutral-toned, slightly bright, balanced body, normally alert
(sexual lust, intoxication altered states of consciousness · normal-paced, relaxed, steady, storytelling) Down below we've got x plus one, x plus three.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, relaxed, steady; timbre is neutral-toned, slightly bright, smooth, balanced body; slurred, frequent disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as sexual lust, intoxication altered states of consciousness; style: storytelling, casual; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 2.1/10; 4.5s, EN.
EN_hj-Jo9BzTa8_W000068 · in -19.5 dBFS · gain -0.5 dB · emolia-02591
(relief · normal-paced, slightly relaxed, moderately variable, casual) And if we leave it in the factored form, it's the easiest way to tell if we've made any mistakes. You can go any farther. So go ahead and take those next two
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as relief; style: casual, playful; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.6/10; 9.4s, EN.
EN_hj-Jo9BzTa8_W000069 · in -19.6 dBFS · gain -0.4 dB · emolia-02591
(normal-paced, slightly relaxed, fairly steady, casual) Subtract those values, simplify as far as you can go.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 1.5/10; 3.5s, EN.
EN_hj-Jo9BzTa8_W000070 · in -21.5 dBFS · gain +1.5 dB · emolia-02591
(pleasure ecstasy · brisk, slightly relaxed, moderately variable, casual) So the first one we've got monomials down below so we can start to build right in the beginning.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as pleasure ecstasy; style: casual, conversational; good recording, quiet background; genuineness 2.2/6; vocal-burst blend 2.3/10; 4.3s, EN.
EN_hj-Jo9BzTa8_W000071 · in -16.2 dBFS · gain -3.8 dB · emolia-02591
(hope enthusiasm optimism, affection, sexual lust · normal-paced, slightly relaxed, fairly steady, casual) And again, grouping together what comes together is going to help you on that second term.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as hope enthusiasm optimism, affection, sexual lust; style: casual, conversational; good recording, no background noise; genuineness 2.7/6; vocal-burst blend 3.4/10; 4.2s, EN.
EN_hj-Jo9BzTa8_W000072 · in -18.8 dBFS · gain -1.2 dB · emolia-02591
Pride(unconstrained axis: Contempt)identity −0.01 emotion 123 %   c-emolia-B1 · #12

This chain comes from the one-sided rule: only Pride had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Pride clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, virtually no clip in this corpus scores higher. That is a total rise of 0.27.

Nothing was asked of the other axis, and in fact Contempt drifts down from 0.92 to 0.84 (-0.08), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.19, then +0.05, then +0.02 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 25 s · fr · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.684 before conversion and 0.673 after — it fell by 0.011. Neighbour-to-neighbour the worst pair went 0.623 → 0.750. (The earlier render, with segment 1 left raw, scores 0.693 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.268 in the original and +0.329 after conversion — 123 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contempt, -0.084 became -0.816.

Quality. Mean predicted overall quality across the segments went 2.79 → 2.99 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.684 → 0.673 -0.011identity cos neighbours 0.623 → 0.750d_b rescored +0.268 → +0.329d_a rescored -0.084 → -0.816d_a mined -0.084d_b mined 0.268min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang frspeaker FR_B00000_S02747total 23.7schain gain +1.4 dBseam step 1.6 dBcrossfades 150/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice
(contempt, anger · fast, energised, slightly tense, dramatic) Pierre ne fit aucun cas des paroles de son grand-père et déclara que les grands garçons n'avaient pas peur des loups.
full caption & clip details
An adult masculine voice; delivery is energised, fast, slightly tense, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; very clear, no disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, guarded; reads as contempt, anger; style: dramatic, cartoonish; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 5.3/10; 6.4s, FR.
FR_B00000_S02747_W000021 · in -16.5 dBFS · gain -3.5 dB · emolia-02645
(pride · measured, normally alert, slightly relaxed, authoritative) Grand-père prit Pierre par le bras et le ramena à la maison.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, rough, thin; clear, no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pride; style: authoritative, dramatic; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 2.8/10; 4.2s, FR.
FR_B00000_S02747_W000022 · in -17.6 dBFS · gain -2.5 dB · emolia-02645
(malevolence malice, contempt, anger · brisk, highly aroused, tense, cartoonish) Le canard se précipita hors de la mare en cacotant.
full caption & clip details
An adult masculine voice; delivery is highly aroused, brisk, tense, volatile; timbre is slightly cool, slightly bright, very rough, thin; very clear, no disfluency, very wide pitch range, audible breath; affect is deeply negative, slightly dominant, guarded; reads as malevolence malice, contempt, anger; style: cartoonish, dramatic; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 3.8/10; 3.8s, FR.
FR_B00000_S02747_W000023 · in -18.8 dBFS · gain -1.2 dB · emolia-02645
(pride, triumph, distress · brisk, highly aroused, tense, dramatic) Mais, malgré tous ses efforts, le loup courrait trop vite. Le voilà qui approche. De plus en plus. Il le rattrape.
full caption & clip details
A middle-aged strongly masculine voice; delivery is highly aroused, brisk, tense, volatile; timbre is slightly cool, very dark, very rough, very full; clear, almost no disfluency, very wide pitch range, audible breath; affect is negative, very dominant, guarded; reads as pride, triumph, distress; style: dramatic, cartoonish; below-average recording, quiet background; genuineness 1.4/6; vocal-burst blend 4.4/10; 9.9s, FR.
FR_B00000_S02747_W000024 · in -18.7 dBFS · gain -1.3 dB · emolia-02645
Confusion(unconstrained axis: Disgust)identity +0.08 emotion 6 %   c-emolia-B1 · #13

This chain comes from the one-sided rule: only Confusion had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Confusion around average — 0.49, lower than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.44.

Nothing was asked of the other axis, and in fact Disgust drifts down from 0.92 to 0.14 (-0.78), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.10, then +0.18, then -0.07, then +0.23 — not a clean run: step 3 moves back the other way by 0.07 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 30 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.636 before conversion and 0.717 after — it rose by 0.081. Neighbour-to-neighbour the worst pair went 0.636 → 0.701. (The earlier render, with segment 1 left raw, scores 0.515 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.436 in the original and +0.026 after conversion — 6 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Disgust, -0.780 became -0.109.

Quality. Mean predicted overall quality across the segments went 2.76 → 3.01 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.636 → 0.717 +0.081identity cos neighbours 0.636 → 0.701d_b rescored +0.436 → +0.026d_a rescored -0.780 → -0.109d_a mined -0.779d_b mined 0.437min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00078_S08253total 28.5schain gain +2.7 dBseam step 0.4 dBcrossfades 100/150/150/100 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, quiet background, normally alert
(disgust · measured, neutral tension, fairly steady, conversational) 如果不是他钓鱼,如果是kiss,全员不在了,嗯,没准儿,我觉得可能有他的比如月迷呀。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust; style: conversational, casual; below-average recording, quiet background; genuineness 4.5/6; vocal-burst blend 3.7/10; 7.9s, ZH.
ZH_B00078_S08253_W000080 · in -17.6 dBFS · gain -2.4 dB · emolia-04055
(amusement, intoxication altered states of consciousness, teasing · measured, relaxed, moderately variable, conversational) (low mumble) 花钱一千多块钱。对嗯,就如果他真能像beatles一样。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, audible breath; affect is mildly negative, slightly submissive, neutral openness; reads as amusement, intoxication altered states of consciousness, teasing; style: conversational, casual; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 1.0/10; 4.8s, ZH.
ZH_B00078_S08253_W000081 · in -19.3 dBFS · gain -0.7 dB · emolia-04055
(doubt, pain · measured, fully relaxed, moderately variable, casual) (low mumble) 嗯嗯,当然宝哥卖卡特你也还在呢嗯嗯。 (ahem)
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, fully relaxed, moderately variable; timbre is neutral-toned, dark, fairly smooth, thin; slurred, frequent disfluency, moderate pitch range, audible breath; affect is mildly negative, neutral stance, neutral openness; style: casual, conversational; poor recording, quiet background; reads as doubt, pain; genuineness 6.0/6; vocal-burst blend 1.4/10; 4.0s, ZH.
ZH_B00078_S08253_W000082 · in -19.9 dBFS · gain -0.1 dB · emolia-04055
(longing · normal-paced, slightly relaxed, fairly steady, conversational) (low mumble) 啊,还还有一点就是我觉得就是刚才想吐槽没吐槽到的,我觉得就是呃打斗戏非常难看。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing; style: conversational, casual; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 5.1/10; 8.2s, ZH.
ZH_B00078_S08253_W000083 · in -18.2 dBFS · gain -1.8 dB · emolia-04055
(confusion · normal-paced, slightly relaxed, moderately variable, conversational) 第二季也难看呀,就就就就一直都挺难看的啊,就是。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; slurred, some disfluency, wide pitch range, audible breath; affect is positive, slightly submissive, neutral openness; reads as confusion; style: conversational, casual; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 3.2/10; 4.3s, ZH.
ZH_B00078_S08253_W000084 · in -18.0 dBFS · gain -2.0 dB · emolia-04055
Emotional Numbness(unconstrained axis: Anger)identity +0.29 emotion 61 %   c-emolia-B1 · #14

This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Emotional Numbness clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.33.

Nothing was asked of the other axis, and in fact Anger drifts down from 0.99 to 0.71 (-0.28), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.10, then +0.15, then +0.07 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.63 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.63 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 29 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.452 before conversion and 0.746 after — it rose by 0.294. Neighbour-to-neighbour the worst pair went 0.452 → 0.746. (The earlier render, with segment 1 left raw, scores 0.588 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.331 in the original and +0.202 after conversion — 61 % of the delta retained. On the other named axis, Anger, -0.284 became -0.188.

Quality. Mean predicted overall quality across the segments went 2.90 → 3.02 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.452 → 0.746 +0.294identity cos neighbours 0.452 → 0.746d_b rescored +0.331 → +0.202d_a rescored -0.284 → -0.188d_a mined -0.284d_b mined 0.331min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00020_S07835total 28.1schain gain +2.2 dBseam step 0.7 dBcrossfades 150/100/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · good recording, slightly relaxed, light breath
(anger, malevolence malice, disgust · measured, normally alert, fairly steady, conversational) And here is actually a dead bird at his feet continue. The man we really must issue a proculamation. The birds are not to be allowed to die here.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, full; very clear, almost no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as anger, malevolence malice, disgust; style: conversational, narration; good recording, quiet background; genuineness 1.0/6; vocal-burst blend 0.1/10; 11.0s, ZH.
ZH_B00020_S07835_W000035 · in -16.7 dBFS · gain -3.3 dB · emolia-03477
(awe · slow, subdued, steady, whispered) So they pulled down the statue of the happy prince.
full caption & clip details
An adult masculine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is slightly warm, slightly dark, rough, very full; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, fairly guarded; reads as awe; style: whispered, monologue; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 4.5/10; 3.3s, ZH.
ZH_B00020_S07835_W000036 · in -25.2 dBFS · gain +5.2 dB · emolia-03477
(contempt, bitterness, sourness · measured, normally alert, moderately variable, narration) As he is no longer beautiful, he is no longer useful, said they are professor for the university.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is slightly warm, neutral-bright, very rough, very full; very clear, almost no disfluency, wide pitch range, light breath; affect is negative, dominant, fairly guarded; reads as contempt, bitterness, sourness; style: narration, storytelling; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 0.7/10; 6.3s, ZH.
ZH_B00020_S07835_W000037 · in -19.0 dBFS · gain -1.0 dB · emolia-03477
(emotional numbness, doubt, malevolence malice · measured, normally alert, fairly steady, narration) Then they melted the statue in a faurness, and the man held a meeting of the corporation to decide what was to be done with the metal.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, doubt, malevolence malice; style: narration, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 8.0s, ZH.
ZH_B00020_S07835_W000038 · in -20.4 dBFS · gain +0.4 dB · emolia-03477
Concentration(unconstrained axis: Emotional Numbness)identity +0.04 emotion 140 %   c-emolia-B1 · #15

This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Concentration clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Emotional Numbness drifts down from 0.84 to 0.13 (-0.71), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.07, then +0.20, then -0.03 — not a clean run: step 3 moves back the other way by 0.03 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.84 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.84 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 32 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.713 before conversion and 0.757 after — it rose by 0.044. Neighbour-to-neighbour the worst pair went 0.794 → 0.761. (The earlier render, with segment 1 left raw, scores 0.632 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.242 in the original and +0.339 after conversion — 140 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, -0.714 became -0.673.

Quality. Mean predicted overall quality across the segments went 2.68 → 2.99 (+0.32) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.713 → 0.757 +0.044identity cos neighbours 0.794 → 0.761d_b rescored +0.242 → +0.339d_a rescored -0.714 → -0.673d_a mined -0.714d_b mined 0.242min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_mPBlCGcGnhgtotal 30.9schain gain +1.4 dBseam step 0.8 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady
(frequent disfluency, casual, didactic) I mean, I guess this is pretty intuitive. But we, we actually walked through this simulator we're going to be using, right?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, didactic; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 2.5/10; 6.3s, EN.
EN_mPBlCGcGnhg_W000008 · in -14.2 dBFS · gain -5.8 dB · emolia-00404
(some disfluency, didactic, formal) But it's, it's a special type of application software. It's, it acts as a simulator. It simulates.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: didactic, formal; good recording, no background noise; genuineness 3.0/6; vocal-burst blend 1.0/10; 5.1s, EN.
EN_mPBlCGcGnhg_W000009 · in -13.2 dBFS · gain -6.8 dB · emolia-00404
(interest, concentration · frequent disfluency, didactic, monologue) Right, so working with QtSpim is synonymous to you using a computer that is built using the MIPS instruction set architecture. So the things we'll be doing with QtSpim.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as interest, concentration; style: didactic, monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.2/10; 11.0s, EN.
EN_mPBlCGcGnhg_W000010 · in -13.8 dBFS · gain -6.2 dB · emolia-00404
(concentration · some disfluency, didactic, monologue) would be running within QtSpim, not on the computer where QtSpim is installed. (ahem) Do you understand this? Right? Similar to the way VirtualBox works.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: didactic, monologue; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 1.2/10; 9.1s, EN.
EN_mPBlCGcGnhg_W000011 · in -15.7 dBFS · gain -4.3 dB · emolia-00404
Concentration(unconstrained axis: Affection)identity +0.11 emotion 39 %   c-emolia-B1 · #16

This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Concentration around average — 0.54, higher than 54 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.37.

Nothing was asked of the other axis, and in fact Affection drifts down from 0.85 to 0.60 (-0.26), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.18, then +0.18, then +0.01 — a plateau around step 3, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.69 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.69 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 32 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.556 before conversion and 0.671 after — it rose by 0.115. Neighbour-to-neighbour the worst pair went 0.734 → 0.700. (The earlier render, with segment 1 left raw, scores 0.484 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.371 in the original and +0.144 after conversion — 39 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Affection, -0.257 became +0.000.

Quality. Mean predicted overall quality across the segments went 2.46 → 2.93 (+0.47) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.556 → 0.671 +0.115identity cos neighbours 0.734 → 0.700d_b rescored +0.371 → +0.144d_a rescored -0.257 → +0.000d_a mined -0.257d_b mined 0.372min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_thndokCgX7ototal 30.9schain gain +2.0 dBseam step 0.8 dBcrossfades 150/100/100 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-bright, quiet background, normal-paced, normally alert, fairly steady
(slightly relaxed, some disfluency, average clarity, casual) And that's what you really need to care about. So being said that,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, playful; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 3.8/10; 3.5s, EN.
EN_thndokCgX7o_W000131 · in -19.6 dBFS · gain -0.4 dB · emolia-02615
(contemplation, shame, helplessness · slightly relaxed, some disfluency, average clarity, casual) This was it. I just want to really share about how, what are the things that you really need to learn it as a tech entrepreneur as well.
full caption & clip details
A young adult somewhat masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, slightly guarded; reads as contemplation, shame, helplessness; style: casual, monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 2.7/10; 6.8s, EN.
EN_thndokCgX7o_W000132 · in -19.4 dBFS · gain -0.6 dB · emolia-02615
(interest, hope enthusiasm optimism, contemplation · slightly relaxed, frequent disfluency, somewhat unclear, casual) What are the things that you really, uh, (low mumble) if you want to build your own startup using the programming technology, if you want to build an app, is it really necessary, is it not, uh, (low mumble) what are the things, what are the channels, what are the psychology, what are the
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, hope enthusiasm optimism, contemplation; style: casual, monologue; below-average recording, quiet background; genuineness 4.0/6; vocal-burst blend 7.1/10; 13.2s, EN.
EN_thndokCgX7o_W000133 · in -18.8 dBFS · gain -1.2 dB · emolia-02615
(concentration, contemplation · relaxed, frequent disfluency, somewhat unclear, casual) various things as well. We'll just in general give you a brief structure. And if you really want to really just kind of maximize the
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as concentration, contemplation; style: casual, monologue; below-average recording, quiet background; genuineness 4.4/6; vocal-burst blend 5.4/10; 7.8s, EN.
EN_thndokCgX7o_W000134 · in -20.6 dBFS · gain +0.6 dB · emolia-02615
Fatigue Exhaustion(unconstrained axis: Pain)identity −0.04 emotion 62 %   c-emolia-B1 · #17

This chain comes from the one-sided rule: only Fatigue Exhaustion had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Fatigue Exhaustion below average — 0.38, lower than 62 % of clips in this corpus — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.45.

Nothing was asked of the other axis, and in fact Pain drifts down from 0.96 to 0.72 (-0.24), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.24, then -0.03 — not a clean run: step 3 moves back the other way by 0.03 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 26 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.865 before conversion and 0.826 after — it fell by 0.039. Neighbour-to-neighbour the worst pair went 0.837 → 0.824. (The earlier render, with segment 1 left raw, scores 0.748 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.450 in the original and +0.277 after conversion — 62 % of the delta retained. On the other named axis, Pain, -0.241 became +0.369.

Quality. Mean predicted overall quality across the segments went 3.11 → 3.18 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.865 → 0.826 -0.039identity cos neighbours 0.837 → 0.824d_b rescored +0.450 → +0.277d_a rescored -0.241 → +0.369d_a mined -0.241d_b mined 0.450min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00035_S00231total 24.7schain gain +1.6 dBseam step 0.9 dBcrossfades 100/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, measured, normally alert
(pain · monologue, narration) 但是林大海根本不听他的劝阻,将林北星推出门去。
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: monologue, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.4/10; 5.2s, ZH.
ZH_B00035_S00231_W000020 · in -18.4 dBFS · gain -1.6 dB · emolia-03623
(doubt, pain · storytelling, dramatic) 空无一人的演出厅里杨朝阳颓废的坐在台边。
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, pain; style: storytelling, dramatic; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 2.2/10; 4.6s, ZH.
ZH_B00035_S00231_W000021 · in -18.4 dBFS · gain -1.6 dB · emolia-03623
(monologue, narration) 面对前来噎郁自己的同学,他根本无法还口,可是高哥却走了进来。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 1.5/10; 6.8s, ZH.
ZH_B00035_S00231_W000022 · in -18.8 dBFS · gain -1.2 dB · emolia-03623
(monologue, narration) 不仅替他狠狠斥责了那两个幸灾乐祸的同学,还在第一排坐下,静静的看起了杨超阳的表演。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.3/10; 8.6s, ZH.
ZH_B00035_S00231_W000023 · in -18.1 dBFS · gain -1.9 dB · emolia-03623
Thankfulness Gratitude(unconstrained axis: Doubt)identity +0.49 emotion 38 %   c-emolia-B1 · #18

This chain comes from the one-sided rule: only Thankfulness Gratitude had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Thankfulness Gratitude around average — 0.55, higher than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.36.

Nothing was asked of the other axis, and in fact Doubt drifts down from 0.96 to 0.90 (-0.06), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.12, then +0.08, then -0.04 — not a clean run: step 4 moves back the other way by 0.04 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of -0.30 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of -0.30 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 59 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.071 before conversion and 0.418 after — it rose by 0.489. Neighbour-to-neighbour the worst pair went 0.052 → 0.247. (The earlier render, with segment 1 left raw, scores 0.337 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.362 in the original and +0.137 after conversion — 38 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Doubt, -0.059 became -0.745.

Quality. Mean predicted overall quality across the segments went 2.83 → 3.07 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 -0.071 → 0.418 +0.489identity cos neighbours 0.052 → 0.247d_b rescored +0.362 → +0.137d_a rescored -0.059 → -0.745d_a mined -0.060d_b mined 0.362min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_veRPNQQ17BQtotal 57.3schain gain +2.0 dBseam step 2.2 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: an elderly masculine voice
(doubt, contentment · normal-paced, very low-energy, neutral tension, casual) So he was telling me about a project that they, you know, they want to do. His vision here is to further educate the children, be able to be self-sufficient, not always depending on, (low mumble) um, let's not call it donations, support, and all that's important and we need that.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as doubt, contentment; style: casual, conversational; average recording, quiet background; genuineness 5.4/6; vocal-burst blend 8.3/10; 15.4s, EN.
EN_veRPNQQ17BQ_W000021 · in -20.1 dBFS · gain +0.1 dB · emolia-00008
(hope enthusiasm optimism, doubt · normal-paced, normally alert, neutral tension, casual) But more for the kids to be able to get to a stage where they're 18, 19 and they're self-sufficient and they can get out in the world and stand on their own feet and, (low mumble) um, earn an honest income.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, slightly thin; somewhat unclear, some disfluency, fairly narrow pitch, audible breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism, doubt; style: casual, monologue; below-average recording, some background noise; genuineness 4.7/6; vocal-burst blend 8.4/10; 9.0s, EN.
EN_veRPNQQ17BQ_W000022 · in -18.4 dBFS · gain -1.6 dB · emolia-00008
(pleasure ecstasy, contentment, affection · brisk, energised, neutral tension, conversational) To see them smile, to see them, you know, open their arms and, and just come and hug us and come and kiss us good morning and just play with them.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, slightly rough, slightly thin; somewhat unclear, some disfluency, wide pitch range, audible breath; affect is positive, slightly submissive, slightly vulnerable; reads as pleasure ecstasy, contentment, affection; style: conversational, monologue; below-average recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.2/10; 9.2s, EN.
EN_veRPNQQ17BQ_W000023 · in -17.5 dBFS · gain -2.5 dB · emolia-00008
(thankfulness gratitude, helplessness, sadness · slow, very low-energy, slightly relaxed, monologue) We're going there (low mumble) to visit family, to be with them, to support them. Whatever one would do for (low mumble) his family, that's (low mumble) what we are trying to do there. Many obligations because of the lack of people on the field from the smallest (low mumble) jobs.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, helplessness, sadness; style: monologue, whispered; below-average recording, quiet background; genuineness 3.7/6; vocal-burst blend 0.7/10; 17.8s, EN.
EN_veRPNQQ17BQ_W000024 · in -17.2 dBFS · gain -2.8 dB · emolia-00008
(thankfulness gratitude, doubt · measured, subdued, relaxed, whispered) May be one of my best uncle or something. And, (low mumble) uh, I really want to thank him.
full caption & clip details
A young adult feminine voice; delivery is subdued, measured, relaxed, steady; timbre is slightly cool, dark, smooth, slightly thin; slurred, frequent disfluency, narrow pitch range, audible breath; affect is neutral, submissive, neutral openness; reads as thankfulness gratitude, doubt; style: whispered, ASMR; below-average recording, quiet background; genuineness 4.5/6; vocal-burst blend 1.7/10; 6.7s, EN.
EN_veRPNQQ17BQ_W000025 · in -17.9 dBFS · gain -2.0 dB · emolia-00008
Contemplation(unconstrained axis: Intoxication Altered States of Consciousness)identity −0.07 emotion 100 %   c-emolia-B1 · #19

This chain comes from the one-sided rule: only Contemplation had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Contemplation below average — 0.37, lower than 63 % of clips in this corpus — and ends with it strongly present at 0.81, higher than 81 % of clips in this corpus. That is a total rise of 0.44.

Nothing was asked of the other axis, and in fact Intoxication Altered States of Consciousness drifts down from 0.82 to 0.71 (-0.11), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.76 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.76 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 14 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.536 before conversion and 0.463 after — it fell by 0.073. Neighbour-to-neighbour the worst pair went 0.594 → 0.463. (The earlier render, with segment 1 left raw, scores 0.515 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.439 in the original and +0.437 after conversion — 100 % of the delta retained, which is essentially all of it. On the other named axis, Intoxication Altered States of Consciousness, -0.107 became +0.050.

Quality. Mean predicted overall quality across the segments went 2.35 → 2.70 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.536 → 0.463 -0.073identity cos neighbours 0.594 → 0.463d_b rescored +0.439 → +0.437d_a rescored -0.107 → +0.050d_a mined -0.107d_b mined 0.439min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_hN__XBF84M4total 13.3schain gain +2.8 dBseam step 0.4 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fairly steady, little disfluency, average clarity, monologue) We've talked about how ecologists study ecosystems.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 1.7/10; 3.0s, EN.
EN_hN__XBF84M4_W000100 · in -17.8 dBFS · gain -2.2 dB · emolia-01649
(steady, little disfluency, clear, monologue) We discovered that ecologists are especially interested in observing the parts of ecosystems.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 0.4/10; 6.1s, EN.
EN_hN__XBF84M4_W000101 · in -19.2 dBFS · gain -0.8 dB · emolia-01649
(fairly steady, some disfluency, average clarity, monologue) The organisms and the non-living parts to draw conclusions about them.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 0.8/10; 4.6s, EN.
EN_hN__XBF84M4_W000102 · in -16.8 dBFS · gain -3.2 dB · emolia-01649
Sexual Lust(unconstrained axis: Fear)identity −0.09 emotion 131 %   c-emolia-B1 · #20

This chain comes from the one-sided rule: only Sexual Lust had to get where it was going, by at least 0.50. The other emotion was left completely free.

The chain starts with Sexual Lust below average — 0.34, lower than 66 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.54.

Nothing was asked of the other axis, and in fact Fear drifts down from 0.77 to 0.65 (-0.12), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.18, then +0.23, then +0.14 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 35 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.873 before conversion and 0.780 after — it fell by 0.093. Neighbour-to-neighbour the worst pair went 0.783 → 0.644. (The earlier render, with segment 1 left raw, scores 0.724 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sexual Lust moved +0.543 in the original and +0.711 after conversion — 131 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Fear, -0.119 became -0.103.

Quality. Mean predicted overall quality across the segments went 2.89 → 2.97 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.873 → 0.780 -0.093identity cos neighbours 0.783 → 0.644d_b rescored +0.543 → +0.711d_a rescored -0.119 → -0.103d_a mined -0.119d_b mined 0.543min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_D6Rl2YztjKQtotal 33.9schain gain +2.4 dBseam step 1.1 dBcrossfades 150/100/100 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed, no disfluency
(normal-paced, steady, light breath, formal) It has also received orders for the 737 AEW&C « Wedgetail » aircraft
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.6s, EN.
EN_D6Rl2YztjKQ_W000226 · in -15.2 dBFS · gain -4.8 dB · emolia-01594
(normal-paced, fairly steady, light breath, formal) The company has also introduced new extended range versions of the 737
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.5/10; 6.6s, EN.
EN_D6Rl2YztjKQ_W000227 · in -15.3 dBFS · gain -4.7 dB · emolia-01594
(triumph · measured, steady, no audible breath, newsreading) The (surprised gasp) 737-900ER is the latest and will extend the range of the 737-900 to a similar range as the successful 737-800 with the capability to fly more passengers.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, no audible breath; affect is neutral, neutral stance, slightly guarded; reads as triumph; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 16.6s, EN.
EN_D6Rl2YztjKQ_W000229 · in -14.8 dBFS · gain -5.2 dB · emolia-01594
(measured, fairly steady, light breath, formal) Due to the addition of two extra emergency exits
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, casual; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 2.0/10; 3.6s, EN.
EN_D6Rl2YztjKQ_W000230 · in -15.3 dBFS · gain -4.7 dB · emolia-01594