This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_c-emolia-PXR.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words
are a real non-speech sound, printed where it happens.
The full generated caption for any clip is under “full caption & clip
details”. Its perceived-gender and background-noise clauses were re-rendered from
the numeric buckets, because the versions stored in the corpus index had those two
ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
The tags and captions describe the ORIGINAL clips — they are the
corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the
converted audio are the ones printed in each card's own paragraph and chips.
Intoxication Altered States of Consciousness ↓ / Emotional Numbness ↑identity −0.18emotion 288 % c-emolia-PXR · #1
This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Emotional Numbness clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.32.
At the same time Intoxication Altered States of Consciousness goes the other way, from 0.70 (higher than 70 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.12 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 18 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.933 before conversion and 0.748 after — it fell by 0.184. Neighbour-to-neighbour the worst pair went 0.949 → 0.793. (The earlier render, with segment 1 left raw, scores 0.625 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.320 in the original and +0.923 after conversion — 288 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Intoxication Altered States of Consciousness, -0.341 became -0.097.
Quality. Mean predicted overall quality across the segments went 2.89 → 2.99 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.933 → 0.748-0.184identity cos neighbours 0.949 → 0.793d_b rescored +0.320 → +0.923d_a rescored -0.341 → -0.097d_a mined -0.341d_b mined 0.319min_cos_consec (site) 0.9466min_cos_anchor (site) 0.9308dataset emolialang enspeaker EN_A0tTdOLv5z0total 17.6schain gain +1.2 dBseam step 0.1 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fairly steady, formal, authoritative)Junzo Yoshimura – 1908–1997 – designed a supplemental building in 1973
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.5/10; 6.4s, EN.
EN_A0tTdOLv5z0_W000003 · in -14.2 dBFS · gain -5.8 dB · emolia-00486
(steady, newsreading, formal)The museum is noted for its collection of Buddhist art, including images, sculpture, and altar articles
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.3/10; 6.0s, EN.
EN_A0tTdOLv5z0_W000004 · in -15.8 dBFS · gain -4.2 dB · emolia-00486
(emotional numbness·fairly steady, formal, monologue)The museum houses and displays works of art belonging to temples and shrines in the Nara area.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.5/10; 5.6s, EN.
EN_A0tTdOLv5z0_W000005 · in -14.9 dBFS · gain -5.1 dB · emolia-00486
This chain comes from the proxy rule: the same two-sided test as above, but because Sourness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Sourness around average — 0.57, higher than 57 % of clips in this corpus — and ends with it at the very top of the corpus at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.33.
At the same time Concentration goes the other way, from 0.92 (higher than 92 % of clips in this corpus) to 0.62 (higher than 62 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are -0.10, then +0.25, then +0.18 — not a clean run: step 1 moves back the other way by 0.10 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 57 s · fr · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.887 before conversion and 0.844 after — it fell by 0.043. Neighbour-to-neighbour the worst pair went 0.851 → 0.869. (The earlier render, with segment 1 left raw, scores 0.740 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sourness moved +0.334 in the original and +0.376 after conversion — 113 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.297 became -0.300.
Quality. Mean predicted overall quality across the segments went 3.00 → 3.29 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.887 → 0.844-0.043identity cos neighbours 0.851 → 0.869d_b rescored +0.334 → +0.376d_a rescored -0.297 → -0.300d_a mined -0.298d_b mined 0.334min_cos_consec (site) 0.8606min_cos_anchor (site) 0.8852dataset emolialang frspeaker FR_Ig0zQk8o_k0total 56.0schain gain +0.9 dBseam step 1.4 dBcrossfades 150/150/100 ms
Script — 4 chunks, 4 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-toned, slightly dark, balanced body, average recording, quiet background, measured, slightly relaxed, fairly steady
(concentration · very low-energy, frequent disfluency, fairly narrow pitch, monologue)où on faisait une ou deux fois par an un culte collectif du quartier et, (low mumble) euh, devant la chapelle du compitum, voilà une autre vue du compitum, du vicus, de la rue d'Asilius.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, didactic; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.4/10; 15.1s, FR.
FR_Ig0zQk8o_k0_W000164 · in -18.4 dBFS · gain -1.6 dB · emolia-02913
(awe·normally alert, frequent disfluency, fairly narrow pitch, monologue)Et, vous voyez, c'est des petites chapelles, (low mumble) euh, celles que je vous ai montrées étaient un peu plus grandes. Et, (ahem) euh, voilà, par exemple, les magistries. Ça, c'est sans doute les quatre magistries.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as awe; style: monologue, didactic; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.7/10; 11.9s, FR.
FR_Ig0zQk8o_k0_W000165 · in -18.1 dBFS · gain -1.9 dB · emolia-02913
(intoxication altered states of consciousness·subdued, frequent disfluency, fairly narrow pitch, monologue)un peu plus relevé, en toge, et puis ces personnages-là qui portent les statuettes des larcs, sans doute, et (low mumble) du génie d'Auguste, (low mumble) euh, étaient les, les esclaves qui formaient un collège avec eux. Et (low mumble) tout ce beau monde fêtait au, vers le début de l'année,
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness; style: monologue; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.6/10; 18.3s, FR.
FR_Ig0zQk8o_k0_W000166 · in -17.8 dBFS · gain -2.2 dB · emolia-02913
(sourness·normally alert, some disfluency, moderate pitch range, monologue)vers le début janvier, hein, (low mumble) fêtait la fête des Compitalia, la fête du Carrefour. Il y avait même des jeux, etc. C'était très prisé pour se faire voir dans le quartier.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as sourness; style: monologue, authoritative; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 1.8/10; 11.2s, FR.
FR_Ig0zQk8o_k0_W000167 · in -19.2 dBFS · gain -0.8 dB · emolia-02913
This chain comes from the proxy rule: the same two-sided test as above, but because Impatience and Irritability is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Impatience and Irritability around average — 0.57, higher than 57 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.28.
At the same time Doubt goes the other way, from 0.74 (higher than 74 % of clips in this corpus) to 0.49 (lower than 51 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.09, then +0.19 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.91 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 31 s · zh · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.852 before conversion and 0.905 after — it rose by 0.053. Neighbour-to-neighbour the worst pair went 0.808 → 0.878. (The earlier render, with segment 1 left raw, scores 0.802 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
The emotional move did not survive. Re-scored end to end, Impatience and Irritability moved +0.279 in the original and -0.044 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Doubt, -0.186 became -0.233.
Quality. Mean predicted overall quality across the segments went 2.96 → 3.19 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.852 → 0.905+0.053identity cos neighbours 0.808 → 0.878d_b rescored +0.279 → -0.044d_a rescored -0.186 → -0.233d_a mined -0.253d_b mined 0.280min_cos_consec (site) 0.9068min_cos_anchor (site) 0.8749dataset emolialang zhspeaker ZH_B00007_S02289total 30.8schain gain -0.4 dBseam step 1.9 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, some disfluency
(measured, average clarity, didactic, authoritative)纳税人为独生子女的,按照每年两万四千元,也就每月两千元的标准定额扣除。纳税人为非独生子女的。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, authoritative; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 3.4/10; 9.6s, ZH.
ZH_B00007_S02289_W000028 · in -19.8 dBFS · gain -0.2 dB · emolia-03351
(normal-paced, average clarity, didactic, monologue)应当与其兄弟姐妹分摊每年两万四千元、每月两千元的扣除额度。分摊方式包括平均分摊、被赡养人指定、分摊或者赡养人约定分摊具体分摊方式在一根纳税年度内不得变更。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 6.3/10; 15.6s, ZH.
ZH_B00007_S02289_W000029 · in -20.3 dBFS · gain +0.3 dB · emolia-03351
This chain comes from the proxy rule: the same two-sided test as above, but because Longing is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Longing clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.27.
At the same time Confusion goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.62 (higher than 62 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.05 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.75 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.82 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.75, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 33 s · zh · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.764 before conversion and 0.724 after — it fell by 0.040. Neighbour-to-neighbour the worst pair went 0.764 → 0.724. (The earlier render, with segment 1 left raw, scores 0.634 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.267 in the original and +0.011 after conversion — 4 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Confusion, -0.334 became -0.258.
Quality. Mean predicted overall quality across the segments went 2.85 → 3.14 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.764 → 0.724-0.040identity cos neighbours 0.764 → 0.724d_b rescored +0.267 → +0.011d_a rescored -0.334 → -0.258d_a mined -0.334d_b mined 0.267min_cos_consec (site) 0.8160min_cos_anchor (site) 0.7466dataset emolialang zhspeaker ZH_B00011_S03574total 32.7schain gain +3.4 dBseam step 2.4 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · neutral-toned, neutral-bright, average recording, quiet background, normally alert, moderately variable, some disfluency
A child feminine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, slightly thin; somewhat unclear, some disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as confusion, astonishment surprise, shame; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.0/6; vocal-burst blend 4.6/10; 10.2s, ZH.
ZH_B00011_S03574_W000050 · in -21.2 dBFS · gain +1.2 dB · emolia-03385
(intoxication altered states of consciousness, fatigue exhaustion, jealousy and envy· measured, neutral tension, average clarity, casual)(ahem) (exhausted groan) (ahem) 就嗯没关系,我下次如果还拿预言家牌,我还要去验这个呃机器猫好吧,我警徽给谁呢?我警徽不知道飞谁,我就不想给你这个,不想给你这个五就贼气,非要掰我的票,这贼气啊不想给这个屋。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as intoxication altered states of consciousness, fatigue exhaustion, jealousy and envy; style: casual, playful; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 6.8/10; 17.6s, ZH.
ZH_B00011_S03574_W000051 · in -18.4 dBFS · gain -1.6 dB · emolia-03385
(longing·normal-paced, slightly relaxed, average clarity, storytelling)你们两个让狼人牌盯着,我看女巫牌走了,还剩一个魔术师一杆枪。
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as longing; style: storytelling, casual; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 3.3/10; 5.3s, ZH.
ZH_B00011_S03574_W000052 · in -17.8 dBFS · gain -2.2 dB · emolia-03385
This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Concentration clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.30.
At the same time Pain goes the other way, from 0.83 (higher than 83 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.06, then +0.24 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 27 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.700 before conversion and 0.585 after — it fell by 0.116. Neighbour-to-neighbour the worst pair went 0.751 → 0.618. (The earlier render, with segment 1 left raw, scores 0.516 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.302 in the original and +0.517 after conversion — 171 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Pain, -0.292 became +0.196.
Quality. Mean predicted overall quality across the segments went 2.65 → 2.83 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.700 → 0.585-0.116identity cos neighbours 0.751 → 0.618d_b rescored +0.302 → +0.517d_a rescored -0.292 → +0.196d_a mined -0.292d_b mined 0.302min_cos_consec (site) 0.8499min_cos_anchor (site) 0.8811dataset emolialang enspeaker EN_vjnGyKqsoAutotal 26.6schain gain +2.5 dBseam step 2.7 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · fairly smooth, quiet background, fairly steady
(fast, normally alert, slightly relaxed, casual)Two affine varieties are isomorphic if and only if their coordinate rings are isomorphic.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 3.3/10; 3.4s, EN.
EN_vjnGyKqsoAu_W000170 · in -20.0 dBFS · gain +0.0 dB · emolia-02477
(normal-paced, normally alert, slightly relaxed, casual)And now this is completely going to be false for projective varieties, okay. So the same
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, didactic; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 1.7/10; 4.6s, EN.
EN_vjnGyKqsoAu_W000171 · in -17.8 dBFS · gain -2.2 dB · emolia-02477
(concentration·measured, subdued, relaxed, monologue)We embedded into different projective spaces and (low mumble) uhhh, if you try to define the uhhh, (ahem) the ring of (low mumble) uhhh, functions (low mumble) uhhh, on that as uhhh, (low mumble) the, this polynomial ring model of the ideal, the homogeneous ideal, you will see that that ring is go, is capable of changing.
full caption & clip details
An elderly masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is slightly cool, dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, narrow pitch range, normal breath; affect is neutral, submissive, slightly guarded; reads as concentration; style: monologue, didactic; below-average recording, quiet background; genuineness 3.8/6; vocal-burst blend 0.9/10; 19.0s, EN.
EN_vjnGyKqsoAu_W000172 · in -20.8 dBFS · gain +0.8 dB · emolia-02477
This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Emotional Numbness around average — 0.51, higher than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.46.
At the same time Sexual Lust goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.21, then +0.05 — a plateau around step 3, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.74 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.77 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.74, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 20 s · zh · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.693 before conversion and 0.761 after — it rose by 0.068. Neighbour-to-neighbour the worst pair went 0.730 → 0.800. (The earlier render, with segment 1 left raw, scores 0.607 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.459 in the original and +0.234 after conversion — 51 % of the delta retained. On the other named axis, Sexual Lust, -0.252 became -0.667.
Quality. Mean predicted overall quality across the segments went 2.55 → 2.74 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.693 → 0.761+0.068identity cos neighbours 0.730 → 0.800d_b rescored +0.459 → +0.234d_a rescored -0.252 → -0.667d_a mined -0.252d_b mined 0.459min_cos_consec (site) 0.7675min_cos_anchor (site) 0.7364dataset emolialang zhspeaker ZH_B00047_S04627total 18.7schain gain +0.4 dBseam step 0.3 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a child feminine voice · dark, thin, very low-energy, frequent disfluency, slurred
This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Infatuation around average — 0.45, lower than 55 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.45.
At the same time Concentration goes the other way, from 0.73 (higher than 73 % of clips in this corpus) to 0.34 (lower than 66 % of clips in this corpus), a change of -0.39. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.06, then +0.14, then +0.24 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.76 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 25 s · zh · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.672 before conversion and 0.630 after — it fell by 0.042. Neighbour-to-neighbour the worst pair went 0.701 → 0.760. (The earlier render, with segment 1 left raw, scores 0.520 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.446 in the original and +0.617 after conversion — 138 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.387 became -0.556.
Quality. Mean predicted overall quality across the segments went 2.88 → 2.99 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.672 → 0.630-0.042identity cos neighbours 0.701 → 0.760d_b rescored +0.446 → +0.617d_a rescored -0.387 → -0.556d_a mined -0.387d_b mined 0.446min_cos_consec (site) 0.7626min_cos_anchor (site) 0.8135dataset emolialang zhspeaker ZH_B00030_S08103total 23.8schain gain +3.3 dBseam step 2.2 dBcrossfades 150/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a middle-aged feminine voice · neutral-toned, fairly smooth, good recording, no background noise, normally alert, slightly relaxed
(measured, moderately variable, some disfluency, playful)八百多篇散文的丰富文化遗产注意啊有两个数字,一千五百四十多首诗歌。八百多篇散文。
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, some disfluency, wide pitch range, audible breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: playful, dramatic; good recording, no background noise; genuineness 2.8/6; vocal-burst blend 2.4/10; 12.7s, ZH.
ZH_B00030_S08103_W000008 · in -19.7 dBFS · gain -0.3 dB · emolia-03575
(impatience and irritability, distress, fear·normal-paced, fairly steady, no disfluency, storytelling)你说他留给我们后人的文章是不是非常多?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as impatience and irritability, distress, fear; style: storytelling, formal; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 4.8/10; 3.9s, ZH.
ZH_B00030_S08103_W000009 · in -18.3 dBFS · gain -1.7 dB · emolia-03575
(confusion·measured, moderately variable, no disfluency, cartoonish)元日,他作于作者初拜相。
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; crisply articulate, no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as confusion; style: cartoonish, dramatic; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 2.6/10; 4.2s, ZH.
ZH_B00030_S08103_W000010 · in -19.7 dBFS · gain -0.3 dB · emolia-03575
(measured, moderately variable, no disfluency, storytelling)正要进行新政改革之时。
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; crisply articulate, no disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; no dominant emotion; style: storytelling, dramatic; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 3.2/10; 3.6s, ZH.
ZH_B00030_S08103_W000011 · in -20.3 dBFS · gain +0.3 dB · emolia-03575
Pain ↓ / Intoxication Altered States of Consciousness ↑identity −0.15emotion 92 % c-emolia-PXR · #8
This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Intoxication Altered States of Consciousness around average — 0.55, higher than 55 % of clips in this corpus — and ends with it strongly present at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.25.
At the same time Pain goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.37. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.08 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 16 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.960 before conversion and 0.806 after — it fell by 0.155. Neighbour-to-neighbour the worst pair went 0.960 → 0.806. (The earlier render, with segment 1 left raw, scores 0.498 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.251 in the original and +0.231 after conversion — 92 % of the delta retained, which is essentially all of it. On the other named axis, Pain, -0.366 became -0.305.
Quality. Mean predicted overall quality across the segments went 2.87 → 2.93 (+0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.960 → 0.806-0.155identity cos neighbours 0.960 → 0.806d_b rescored +0.251 → +0.231d_a rescored -0.366 → -0.305d_a mined -0.366d_b mined 0.251min_cos_consec (site) 0.9595min_cos_anchor (site) 0.9587dataset emolialang enspeaker EN_B00061_S08450total 14.8schain gain +1.6 dBseam step 0.8 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(pain · fairly steady, formal, monologue)On March 2, 2002, an incident happened, hashid or hamafutsal
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.8/10; 5.4s, EN.
EN_B00061_S08450_W000001 · in -16.6 dBFS · gain -3.5 dB · emolia-01407
(steady, formal, authoritative)In 2005, Weatherman Danny Roop announced that he moved to Channel 10
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.0/10; 5.3s, EN.
EN_B00061_S08450_W000002 · in -14.6 dBFS · gain -5.4 dB · emolia-01407
(fairly steady, formal, authoritative)This caused the news company's deficit to grow to 10 million shekels
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.8/10; 4.6s, EN.
EN_B00061_S08450_W000003 · in -14.3 dBFS · gain -5.7 dB · emolia-01407
Intoxication Altered States of Consciousness ↓ / Affection ↑identity +0.03emotion 48 % c-emolia-PXR · #9
This chain comes from the proxy rule: the same two-sided test as above, but because Affection is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Affection clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.36.
At the same time Intoxication Altered States of Consciousness goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.44. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.16, then +0.01, then +0.19 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.71 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.74 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.71, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 22 s · zh · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.711 before conversion and 0.736 after — it rose by 0.025. Neighbour-to-neighbour the worst pair went 0.744 → 0.757. (The earlier render, with segment 1 left raw, scores 0.681 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Affection moved +0.364 in the original and +0.174 after conversion — 48 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Intoxication Altered States of Consciousness, -0.443 became -0.285.
Quality. Mean predicted overall quality across the segments went 2.97 → 3.03 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.711 → 0.736+0.025identity cos neighbours 0.744 → 0.757d_b rescored +0.364 → +0.174d_a rescored -0.443 → -0.285d_a mined -0.443d_b mined 0.364min_cos_consec (site) 0.7435min_cos_anchor (site) 0.7103dataset emolialang zhspeaker ZH_B00034_S00702total 21.1schain gain +2.6 dBseam step 1.7 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a child feminine voice · balanced body, no background noise
(intoxication altered states of consciousness, pain, teasing · measured, normally alert, relaxed, storytelling)一人中途而废,另一人也做不得事。
full caption & clip details
A child feminine voice; delivery is normally alert, measured, relaxed, moderately variable; timbre is slightly warm, dark, slightly rough, balanced body; slurred, frequent disfluency, moderate pitch range, audible breath; affect is mildly positive, slightly submissive, neutral openness; reads as intoxication altered states of consciousness, pain, teasing; style: storytelling, whispered; average recording, no background noise; genuineness 2.8/6; vocal-burst blend 3.2/10; 5.7s, ZH.
ZH_B00034_S00702_W000002 · in -18.0 dBFS · gain -2.0 dB · emolia-03614
(doubt, intoxication altered states of consciousness, teasing ·slow, very low-energy, slightly relaxed, ASMR)俗话说,一刻不凡二主这事儿。
full caption & clip details
A child feminine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; crisply articulate, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; reads as doubt, intoxication altered states of consciousness, teasing; style: ASMR, whispered; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.2/10; 5.4s, ZH.
ZH_B00034_S00702_W000003 · in -18.4 dBFS · gain -1.6 dB · emolia-03614
This chain comes from the proxy rule: the same two-sided test as above, but because Longing is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Longing clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.37.
At the same time Pleasure Ecstasy goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.74 (higher than 74 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.22 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.73 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.74 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.73, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 38 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.724 before conversion and 0.772 after — it rose by 0.049. Neighbour-to-neighbour the worst pair went 0.773 → 0.800. (The earlier render, with segment 1 left raw, scores 0.654 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.372 in the original and +0.439 after conversion — 118 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Pleasure Ecstasy, -0.254 became -0.311.
Quality. Mean predicted overall quality across the segments went 2.80 → 3.07 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.724 → 0.772+0.049identity cos neighbours 0.773 → 0.800d_b rescored +0.372 → +0.439d_a rescored -0.254 → -0.311d_a mined -0.254d_b mined 0.372min_cos_consec (site) 0.7352min_cos_anchor (site) 0.7267dataset emolialang enspeaker EN_gQWG2De90Sototal 37.2schain gain +1.8 dBseam step 1.4 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · slightly cool
(interest, pleasure ecstasy, hope enthusiasm optimism · fast, highly aroused, tense, dramatic)It's buried in good things called ministry. It's buried in good things like writing a book, the next travel gig, the next thing, the healing testimonies, all of this stuff. He said your face is glued to everything. But he said remember the day where you had nothing to look at except my face.
full caption & clip details
A young adult masculine voice; delivery is highly aroused, fast, tense, volatile; timbre is slightly cool, slightly bright, very rough, thin; clear, some disfluency, very wide pitch range, audible breath; affect is elated, very dominant, guarded; reads as interest, pleasure ecstasy, hope enthusiasm optimism; style: dramatic, authoritative; below-average recording, noisy background; genuineness 3.2/6; vocal-burst blend 5.4/10; 16.7s, EN.
EN_gQWG2De90So_W000105 · in -15.2 dBFS · gain -4.8 dB · emolia-02550
(anger, disappointment, distress·brisk, highly aroused, tense, dramatic)Your back was toward everything else and it seems like blessings started to find you when you were not looking for them and now your face turned toward those blessings and you might not have realized but you actually (ahem) turned your back toward my face.
full caption & clip details
A young adult masculine voice; delivery is highly aroused, brisk, tense, variable; timbre is slightly cool, slightly bright, very rough, thin; very clear, almost no disfluency, very wide pitch range, normal breath; affect is elated, very dominant, guarded; reads as anger, disappointment, distress; style: dramatic, ranting; below-average recording, some background noise; genuineness 3.0/6; vocal-burst blend 4.3/10; 12.1s, EN.
EN_gQWG2De90So_W000106 · in -16.2 dBFS · gain -3.8 dB · emolia-02550
(longing, affection, infatuation·measured, energised, slightly tense, dramatic)The Lord says to some of you here today, I wanna see your face. I see your body here, but I wanna see your face.
full caption & clip details
A middle-aged strongly masculine voice; delivery is energised, measured, slightly tense, variable; timbre is slightly cool, neutral-bright, rough, full; very clear, almost no disfluency, wide pitch range, audible breath; affect is negative, slightly dominant, vulnerable; reads as longing, affection, infatuation; style: dramatic, storytelling; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 1.4/10; 8.8s, EN.
EN_gQWG2De90So_W000107 · in -19.2 dBFS · gain -0.8 dB · emolia-02550
Intoxication Altered States of Consciousness ↓ / Bitterness ↑identity −0.09emotion REVERSED c-emolia-PXR · #11
This chain comes from the proxy rule: the same two-sided test as above, but because Bitterness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Bitterness clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.31.
At the same time Intoxication Altered States of Consciousness goes the other way, from 0.80 (higher than 80 % of clips in this corpus) to 0.51 (right about the corpus median), a change of -0.29. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.15 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 20 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.856 before conversion and 0.765 after — it fell by 0.091. Neighbour-to-neighbour the worst pair went 0.888 → 0.766. (The earlier render, with segment 1 left raw, scores 0.675 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
The emotional move did not survive. Re-scored end to end, Bitterness moved +0.308 in the original and -0.168 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Intoxication Altered States of Consciousness, -0.295 became -0.379.
Quality. Mean predicted overall quality across the segments went 2.74 → 2.90 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.856 → 0.765-0.091identity cos neighbours 0.888 → 0.766d_b rescored +0.308 → -0.168d_a rescored -0.295 → -0.379d_a mined -0.295d_b mined 0.308min_cos_consec (site) 0.8901min_cos_anchor (site) 0.8888dataset emolialang enspeaker EN_M7aCInQUHsUtotal 19.3schain gain +1.6 dBseam step 0.4 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · fairly smooth, balanced body, average recording, quiet background, fairly steady, moderate pitch range
(normal-paced, normally alert, slightly relaxed, casual)That's what satisfied means, right? To do something to satisfaction.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.5/10; 4.1s, EN.
EN_M7aCInQUHsU_W000402 · in -25.4 dBFS · gain +5.4 dB · emolia-01240
(pain· normal-paced, normally alert, slightly relaxed, casual)(ahem) Uh, and he's unable to describe their kindness to his satisfaction because they are simply too kind.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: casual, monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.4/10; 7.4s, EN.
EN_M7aCInQUHsU_W000403 · in -23.0 dBFS · gain +3.0 dB · emolia-01240
(bitterness, malevolence malice, contempt·slow, very low-energy, relaxed, monologue)Especially about Mrs. Harville taking care of Louisa. Exertions means efforts.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly cool, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, submissive, neutral openness; reads as bitterness, malevolence malice, contempt; style: monologue, didactic; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 0.0/10; 8.3s, EN.
EN_M7aCInQUHsU_W000404 · in -23.1 dBFS · gain +3.1 dB · emolia-01240
This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Contemplation below average — 0.35, lower than 65 % of clips in this corpus — and ends with it strongly present at 0.78, higher than 78 % of clips in this corpus. That is a total rise of 0.43.
At the same time Concentration goes the other way, from 0.82 (higher than 82 % of clips in this corpus) to 0.33 (lower than 67 % of clips in this corpus), a change of -0.50. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.04, then +0.23, then +0.17 — a plateau around step 1, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.91 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 29 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.805 before conversion and 0.767 after — it fell by 0.038. Neighbour-to-neighbour the worst pair went 0.880 → 0.804. (The earlier render, with segment 1 left raw, scores 0.558 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.433 in the original and +0.646 after conversion — 149 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.498 became -0.529.
Quality. Mean predicted overall quality across the segments went 2.94 → 3.06 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.805 → 0.767-0.038identity cos neighbours 0.880 → 0.804d_b rescored +0.433 → +0.646d_a rescored -0.498 → -0.529d_a mined -0.498d_b mined 0.432min_cos_consec (site) 0.9116min_cos_anchor (site) 0.9229dataset emolialang enspeaker EN_T4fdk8KgEZ4total 27.9schain gain +1.3 dBseam step 1.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fairly steady, formal, newsreading)Pollinin core samples from Lake Baffa in the Latmas region inland of Miletus suggests that a lightly grazed climax forest prevailed in the Meander Valley, otherwise untenanted
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 9.8s, EN.
EN_T4fdk8KgEZ4_W000027 · in -14.2 dBFS · gain -5.8 dB · emolia-01725
(fairly steady, formal, newsreading)Sparse Neolithic settlements were made at springs, numerous and sometimes geothermal in this karst, rift valley topography
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.2/10; 7.1s, EN.
EN_T4fdk8KgEZ4_W000028 · in -14.8 dBFS · gain -5.2 dB · emolia-01725
(emotional numbness, infatuation· fairly steady, formal, authoritative)The islands offshore were settled perhaps for their strategic significance at the mouth of the Meander, a route inland protected by escarpments
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, infatuation; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 7.4s, EN.
EN_T4fdk8KgEZ4_W000029 · in -15.3 dBFS · gain -4.7 dB · emolia-01725
(steady, formal, monologue)The graziers in the valley may have belonged to them, but the location looked to the sea
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 1.3/10; 4.2s, EN.
EN_T4fdk8KgEZ4_W000030 · in -14.8 dBFS · gain -5.2 dB · emolia-01725
This chain comes from the proxy rule: the same two-sided test as above, but because Fear is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Fear clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.28.
At the same time Longing goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.11 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.88 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 42 s · zh · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.833 before conversion and 0.846 after — it rose by 0.013. Neighbour-to-neighbour the worst pair went 0.759 → 0.792. (The earlier render, with segment 1 left raw, scores 0.784 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.283 in the original and +0.267 after conversion — 94 % of the delta retained, which is essentially all of it. On the other named axis, Longing, -0.255 became -0.082.
Quality. Mean predicted overall quality across the segments went 2.96 → 2.97 (+0.01) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.833 → 0.846+0.013identity cos neighbours 0.759 → 0.792d_b rescored +0.283 → +0.267d_a rescored -0.255 → -0.082d_a mined -0.255d_b mined 0.283min_cos_consec (site) 0.8800min_cos_anchor (site) 0.9051dataset emolialang zhspeaker ZH_B00067_S02048total 41.3schain gain +2.4 dBseam step 2.0 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-bright, fairly smooth, good recording, no background noise, clear
(longing, embarrassment, astonishment surprise · measured, very low-energy, slightly relaxed, ASMR)I didn't want to say that i'd never heard of vizabella rosli. I'm going to meet the camera crew at the law courts.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, vulnerable; reads as longing, embarrassment, astonishment surprise; style: ASMR, whispered; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 3.2/10; 7.0s, ZH.
ZH_B00067_S02048_W000533 · in -16.9 dBFS · gain -3.1 dB · emolia-03950
(doubt, emotional numbness, fatigue exhaustion·normal-paced, normally alert, slightly relaxed, narration)I have to report on a story for television without knowing what it is about.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as doubt, emotional numbness, fatigue exhaustion; style: narration, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.8/10; 4.9s, ZH.
ZH_B00067_S02048_W000534 · in -16.8 dBFS · gain -3.2 dB · emolia-03950
(fear, doubt, teasing· normal-paced, energised, neutral tension, conversational)I just came out of the toilets and found paculi outside with richards. Dogs. Are you? Okay? She asked, you look a bit worried. No, no, i'm fine. I said. Are you sure she stared at me for a moment? Listen, richard didn't mean, is a bella roselinhe 's thinking of elana rocini right ah now i understand.
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, neutral tension, variable; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; clear, some disfluency, wide pitch range, normal breath; affect is negative, slightly dominant, neutral openness; reads as fear, doubt, teasing; style: conversational, storytelling; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 3.5/10; 29.7s, ZH.
ZH_B00067_S02048_W000535 · in -17.5 dBFS · gain -2.5 dB · emolia-03950
This chain comes from the proxy rule: the same two-sided test as above, but because Longing is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Longing clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.28.
At the same time Contempt goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.12, then +0.17, then -0.19, then +0.19 — not a clean run: step 3 moves back the other way by 0.19 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 74 s · zh · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.774 before conversion and 0.674 after — it fell by 0.100. Neighbour-to-neighbour the worst pair went 0.774 → 0.674. (The earlier render, with segment 1 left raw, scores 0.656 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.278 in the original and +0.197 after conversion — 71 % of the delta retained, which is most of it. On the other named axis, Contempt, -0.329 became -0.127.
Quality. Mean predicted overall quality across the segments went 2.68 → 3.18 (+0.50) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.774 → 0.674-0.100identity cos neighbours 0.774 → 0.674d_b rescored +0.278 → +0.197d_a rescored -0.329 → -0.127d_a mined -0.329d_b mined 0.277min_cos_consec (site) 0.8846min_cos_anchor (site) 0.8980dataset emolialang zhspeaker ZH_26k0Hylr1qototal 73.1schain gain +4.3 dBseam step 0.9 dBcrossfades 100/150/100/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, normally alert, moderately variable, some disfluency, light breath
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as elation, triumph, interest; style: casual, conversational; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 8.8/10; 21.2s, ZH.
ZH_26k0Hylr1qo_W000042 · in -21.7 dBFS · gain +1.7 dB · emolia-03276
(jealousy and envy, longing, infatuation· brisk, neutral tension, average clarity, casual)咁呢度佢有櫃啦,佢就冇門再通去洗手間啦,佢呢度係open嘅,咁我見好多比較外國式嘅住宅,其實都鍾意唔再裝門。我入咗嚟呢個位呢,其實已經係shower嘅位置㗎啦,咁就可以30啦,咁呢個frameless嘅位呢,咁呢度係其實又有一級嘅,所以啲水呢應該係唔會脹出嚟嘅。
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as jealousy and envy, longing, infatuation; style: casual, dramatic; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 9.9/10; 19.3s, ZH.
ZH_26k0Hylr1qo_W000043 · in -19.7 dBFS · gain -0.3 dB · emolia-03276
(teasing, sexual lust, elation·normal-paced, slightly relaxed, average clarity, dramatic)咁呢度全個呢個位置呢都係鋪咗磚啦咁呢度有花灑咁有水龍頭然後浴缸呢就喺入面嘅咁長身呢呢邊係玻璃啦呢度就全部都係鋪咗磚嘅shower去到浴缸呢一個位呢都有一級嘅
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as teasing, sexual lust, elation; style: dramatic, playful; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 6.2/10; 16.7s, ZH.
ZH_26k0Hylr1qo_W000044 · in -19.8 dBFS · gain -0.2 dB · emolia-03276
(longing, relief, thankfulness gratitude· normal-paced, neutral tension, average clarity, monologue)咁所以佢就確保大家可能係淋浴嘅時候啲水其實都會喺呢個低位範圍啦咁佢個排水呢就係呢一條喇掛毛巾位喺後面
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing, relief, thankfulness gratitude; style: monologue, casual; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 3.0/10; 10.6s, ZH.
ZH_26k0Hylr1qo_W000045 · in -21.1 dBFS · gain +1.1 dB · emolia-03276
This chain comes from the proxy rule: the same two-sided test as above, but because Contentment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Contentment clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.26.
At the same time Malevolence Malice goes the other way, from 0.74 (higher than 74 % of clips in this corpus) to 0.22 (lower than 78 % of clips in this corpus), a change of -0.52. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.14, then -0.11, then +0.01, then +0.21 — not a clean run: step 2 moves back the other way by 0.11 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 59 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.823 before conversion and 0.813 after — it fell by 0.010. Neighbour-to-neighbour the worst pair went 0.852 → 0.829. (The earlier render, with segment 1 left raw, scores 0.745 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.256 in the original and +0.416 after conversion — 162 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Malevolence Malice, -0.521 became -0.366.
Quality. Mean predicted overall quality across the segments went 2.94 → 3.07 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.823 → 0.813-0.010identity cos neighbours 0.852 → 0.829d_b rescored +0.256 → +0.416d_a rescored -0.521 → -0.366d_a mined -0.521d_b mined 0.257min_cos_consec (site) 0.8525min_cos_anchor (site) 0.8533dataset emolialang enspeaker EN_xmwk7bBd8QAtotal 57.5schain gain +2.5 dBseam step 1.2 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, good recording, normal-paced, normally alert, some disfluency
(slightly relaxed, fairly steady, average clarity, casual)While the student is recording their video. Now this is kind of an interesting little control area of your grid settings.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, authoritative; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 0.1/10; 8.2s, EN.
EN_xmwk7bBd8QA_W000029 · in -22.2 dBFS · gain +2.2 dB · emolia-01296
(relief, distress, fear·relaxed, moderately variable, average clarity, casual)Say that you're covering a sensitive topic and you're worried that a student might create an inappropriate, thoughtless, or rude video that everyone would see and it would just be a headache.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as relief, distress, fear; style: casual, conversational; good recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.7/10; 12.4s, EN.
EN_xmwk7bBd8QA_W000030 · in -25.1 dBFS · gain +5.1 dB · emolia-01296
(slightly relaxed, moderately variable, average clarity, casual)You can choose to have your videos moderated. And what that means is if you turn this on, as soon as the student submits a video, it'll let you know.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 4.0/10; 7.8s, EN.
EN_xmwk7bBd8QA_W000031 · in -21.9 dBFS · gain +1.9 dB · emolia-01296
(sourness, contempt, infatuation· slightly relaxed, fairly steady, clear, casual)You can go in, view the video, and if it meets the guidelines that you're wanting and is appropriate, then you can go ahead and publish it to the grid. But nothing will automatically be published without you taking a look at it. This makes a little bit more work for you, but it also helps avoid potential headaches.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, full; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as sourness, contempt, infatuation; style: casual, conversational; good recording, no background noise; mildly explicit content; genuineness 1.2/6; vocal-burst blend 2.7/10; 17.5s, EN.
EN_xmwk7bBd8QA_W000032 · in -25.5 dBFS · gain +5.5 dB · emolia-01296
(contentment· slightly relaxed, fairly steady, average clarity, casual)You can also allow students to reply to each other. I recommend turning this on, (low mumble) uhm, because it allows for more of that interaction and feedback rather than just passively watching the videos.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contentment; style: casual, whispered; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 1.3/10; 12.5s, EN.
EN_xmwk7bBd8QA_W000033 · in -21.6 dBFS · gain +1.6 dB · emolia-01296
This chain comes from the proxy rule: the same two-sided test as above, but because Disgust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Disgust clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.34.
At the same time Contemplation goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.10 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.74 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.83 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.74, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 40 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.727 before conversion and 0.731 after — it rose by 0.004. Neighbour-to-neighbour the worst pair went 0.727 → 0.731. (The earlier render, with segment 1 left raw, scores 0.730 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disgust moved +0.342 in the original and +0.118 after conversion — 35 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Contemplation, -0.269 became -0.223.
Quality. Mean predicted overall quality across the segments went 2.75 → 3.12 (+0.37) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.727 → 0.731+0.004identity cos neighbours 0.727 → 0.731d_b rescored +0.342 → +0.118d_a rescored -0.269 → -0.223d_a mined -0.269d_b mined 0.342min_cos_consec (site) 0.8282min_cos_anchor (site) 0.7409dataset emolialang enspeaker EN_bz3KsFrmE0utotal 38.9schain gain +3.2 dBseam step 0.9 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, full, good recording, no background noise, measured, clear
(contemplation, anger, bitterness · energised, slightly relaxed, fairly steady, dramatic)They also think and feel that if they truly do have everything in life and what you're saying is true.
full caption & clip details
An adult masculine voice; delivery is energised, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, fairly guarded; reads as contemplation, anger, bitterness; style: dramatic, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.1/10; 7.2s, EN.
EN_bz3KsFrmE0u_W000019 · in -23.2 dBFS · gain +3.2 dB · emolia-01283
(bitterness, fear, impatience and irritability· energised, neutral tension, fairly steady, dramatic)But they still feel this way and they still feel depressed. Then nothing's ever going to get better, is it? If you could ignore your depression, you would, wouldn't you? But it's not that simple and it's something that you just cannot do.
full caption & clip details
An adult masculine voice; delivery is energised, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, full; clear, some disfluency, wide pitch range, normal breath; affect is negative, slightly dominant, fairly guarded; reads as bitterness, fear, impatience and irritability; style: dramatic, storytelling; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 3.1/10; 19.3s, EN.
EN_bz3KsFrmE0u_W000020 · in -23.9 dBFS · gain +4.0 dB · emolia-01283
(contempt, disgust, malevolence malice·normally alert, slightly relaxed, steady, whispered)Saying this to somebody who is depressed is a pretty stupid thing to say. So please do not say this. Number six, you're not the same and you're not fun to be around anymore.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, slightly rough, full; clear, little disfluency, moderate pitch range, normal breath; affect is mildly negative, slightly dominant, fairly guarded; reads as contempt, disgust, malevolence malice; style: whispered, didactic; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.7/10; 12.8s, EN.
EN_bz3KsFrmE0u_W000021 · in -23.6 dBFS · gain +3.6 dB · emolia-01283
This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Emotional Numbness clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.30.
At the same time Jealousy and Envy goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.07, then +0.22 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 53 s · de · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.772 before conversion and 0.768 after — it fell by 0.004. Neighbour-to-neighbour the worst pair went 0.909 → 0.854. (The earlier render, with segment 1 left raw, scores 0.608 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.297 in the original and +0.282 after conversion — 95 % of the delta retained, which is essentially all of it. On the other named axis, Jealousy and Envy, -0.313 became -0.653.
Quality. Mean predicted overall quality across the segments went 3.06 → 3.33 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.772 → 0.768-0.004identity cos neighbours 0.909 → 0.854d_b rescored +0.297 → +0.282d_a rescored -0.313 → -0.653d_a mined -0.313d_b mined 0.297min_cos_consec (site) 0.9541min_cos_anchor (site) 0.9467dataset emolialang despeaker DE_8d28v27HEWItotal 52.6schain gain -0.1 dBseam step 2.1 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert, slightly relaxed, fairly steady
(jealousy and envy, impatience and irritability, teasing · normal-paced, monologue, casual)Ungroup.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, impatience and irritability, teasing; style: monologue, casual; good recording, quiet background; genuineness 2.5/6; vocal-burst blend 5.4/10; 30.0s, DE.
DE_8d28v27HEWI_W000036 · in -15.9 dBFS · gain -4.0 dB · emolia-00024
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, concentration; style: monologue, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 3.1/10; 15.2s, DE.
DE_8d28v27HEWI_W000037 · in -16.5 dBFS · gain -3.5 dB · emolia-00024
(emotional numbness, intoxication altered states of consciousness·fast, casual)select, move, resize, and rotate, und so weiter. Das bringt euch aber alles noch nicht, solange ihr kein Terrain habt. Und dafür gehen wir einfach mal hier oben auf edit.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness, intoxication altered states of consciousness; style: casual; good recording, quiet background; genuineness 3.5/6; vocal-burst blend 5.8/10; 7.7s, DE.
DE_8d28v27HEWI_W000038 · in -17.0 dBFS · gain -3.0 dB · emolia-00024
This chain comes from the proxy rule: the same two-sided test as above, but because Hope Enthusiasm Optimism is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Hope Enthusiasm Optimism around average — 0.57, higher than 57 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.39.
At the same time Impatience and Irritability goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.42 (lower than 58 % of clips in this corpus), a change of -0.52. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.21, then -0.12, then +0.13 — not a clean run: step 3 moves back the other way by 0.12 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 73 s · en · emolia
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.771 before conversion and 0.775 after — it rose by 0.004. Neighbour-to-neighbour the worst pair went 0.771 → 0.775. (The earlier render, with segment 1 left raw, scores 0.688 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.387 in the original and +0.363 after conversion — 94 % of the delta retained, which is essentially all of it. On the other named axis, Impatience and Irritability, -0.517 became -0.313.
Quality. Mean predicted overall quality across the segments went 3.04 → 3.24 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.771 → 0.775+0.004identity cos neighbours 0.771 → 0.775d_b rescored +0.387 → +0.363d_a rescored -0.517 → -0.313d_a mined -0.517d_b mined 0.387min_cos_consec (site) 0.8551min_cos_anchor (site) 0.8467dataset emolialang enspeaker EN_B00061_S06045total 71.5schain gain +1.6 dBseam step 1.3 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, quiet background, normally alert, some disfluency, light breath
(impatience and irritability, bitterness · normal-paced, neutral tension, moderately variable, didactic)But let me say this, some of my best traders only trade two or three setups. It's all they trade with. So there are literally, (low mumble) uh, well,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as impatience and irritability, bitterness; style: didactic, casual; good recording, quiet background; genuineness 3.2/6; vocal-burst blend 3.2/10; 9.7s, EN.
EN_B00061_S06045_W000006 · in -20.2 dBFS · gain +0.2 dB · emolia-01416
(brisk, neutral tension, moderately variable, casual)In the new members area that's about to be launched, (ahem) uh, you'll click on there and you'll receive a ticket number. So if you've got a question, (low mumble) uh, I'm actually hiring additional coaches and moderators, which I'll explain in a moment, you'll have a ticket number and I've got to then respond back within a certain amount of time. So I can just track quality control if you like. Live trading room we'll talk about in a moment, members only, swing trading, futures trading, markets to trade.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 4.8/10; 25.3s, EN.
EN_B00061_S06045_W000007 · in -26.6 dBFS · gain +6.6 dB · emolia-01416
(sourness, hope enthusiasm optimism, pride·normal-paced, slightly relaxed, fairly steady, didactic)This is my philosophy. So rather than spend a fortune in advertising and marketing selling expensive training and coaching programs, to hopefully have some of those traders join my room,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, hope enthusiasm optimism, pride; style: didactic, monologue; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 1.0/10; 11.3s, EN.
EN_B00061_S06045_W000008 · in -21.8 dBFS · gain +1.8 dB · emolia-01416
(affection, amusement· normal-paced, slightly relaxed, fairly steady, casual)And I've got to tell you this, my wife calls and love letters, which I sort of like. I receive emails from members.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as affection, amusement; style: casual, playful; good recording, quiet background; genuineness 4.1/6; vocal-burst blend 1.6/10; 6.6s, EN.
EN_B00061_S06045_W000009 · in -23.8 dBFS · gain +3.8 dB · emolia-01416
(hope enthusiasm optimism, interest·brisk, slightly relaxed, moderately variable, casual)Now we've got some great counter trend strategies by the way, which I'll quickly show you in a moment. Uh, (low mumble) but really start one thing and develop over a period of time. Now, if you're an experienced trader, uh, (ahem) you've got to go back and draw a line in the sand and basically restart. And that's something I talk a lot about in my program.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, interest; style: casual, conversational; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 4.3/10; 19.5s, EN.
EN_B00061_S06045_W000010 · in -23.5 dBFS · gain +3.5 dB · emolia-01416
This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Contemplation clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.26.
At the same time Fatigue Exhaustion goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.01 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.84 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 23 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.643 before conversion and 0.608 after — it fell by 0.035. Neighbour-to-neighbour the worst pair went 0.742 → 0.695. (The earlier render, with segment 1 left raw, scores 0.426 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.260 in the original and +0.080 after conversion — 31 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Fatigue Exhaustion, -0.265 became -0.094.
Quality. Mean predicted overall quality across the segments went 2.67 → 2.87 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.643 → 0.608-0.035identity cos neighbours 0.742 → 0.695d_b rescored +0.260 → +0.080d_a rescored -0.265 → -0.094d_a mined -0.265d_b mined 0.259min_cos_consec (site) 0.8366min_cos_anchor (site) 0.7897dataset emolialang enspeaker EN_Wq9ts857dSutotal 22.6schain gain +2.0 dBseam step 0.7 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · slightly warm, dark, smooth, slow, very low-energy, relaxed, steady, frequent disfluency
A young adult masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, smooth, very full; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly negative, submissive, neutral openness; reads as fatigue exhaustion, emotional numbness, shame; style: whispered, monologue; average recording, no background noise; genuineness 2.8/6; vocal-burst blend 6.1/10; 3.2s, EN.
EN_Wq9ts857dSu_W000013 · in -20.1 dBFS · gain +0.1 dB · emolia-02475
(contemplation, helplessness, sadness·slurred, narrow pitch range, normal breath, whispered)Some people have such problems that they seem incapable of practicing.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, smooth, very full; slurred, frequent disfluency, narrow pitch range, normal breath; affect is mildly positive, submissive, neutral openness; reads as contemplation, helplessness, sadness; style: whispered, ASMR; below-average recording, quiet background; genuineness 1.5/6; vocal-burst blend 4.6/10; 8.8s, EN.
EN_Wq9ts857dSu_W000014 · in -19.1 dBFS · gain -0.9 dB · emolia-02475
(contemplation, concentration, doubt· slurred, narrow pitch range, normal breath, whispered)We're not just talking about struggle. We're talking about the incapacity to practice meditation. It could be because of a neurotic impulse or
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, smooth, balanced body; slurred, frequent disfluency, narrow pitch range, normal breath; affect is mildly negative, submissive, neutral openness; reads as contemplation, concentration, doubt; style: whispered, ASMR; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 4.4/10; 11.0s, EN.
EN_Wq9ts857dSu_W000015 · in -20.1 dBFS · gain +0.1 dB · emolia-02475
Intoxication Altered States of Consciousness ↓ / Emotional Numbness ↑identity −0.06emotion 62 % c-emolia-PXR · #20
This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Emotional Numbness around average — 0.49, right about the corpus median — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.42.
At the same time Intoxication Altered States of Consciousness goes the other way, from 0.74 (higher than 74 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.23 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.78 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.78 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.78, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 21 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.726 before conversion and 0.662 after — it fell by 0.064. Neighbour-to-neighbour the worst pair went 0.768 → 0.741. (The earlier render, with segment 1 left raw, scores 0.504 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.410 in the original and +0.255 after conversion — 62 % of the delta retained. On the other named axis, Intoxication Altered States of Consciousness, -0.375 became -0.353.
Quality. Mean predicted overall quality across the segments went 2.55 → 2.79 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.726 → 0.662-0.064identity cos neighbours 0.768 → 0.741d_b rescored +0.410 → +0.255d_a rescored -0.375 → -0.353d_a mined -0.375d_b mined 0.423min_cos_consec (site) 0.7843min_cos_anchor (site) 0.7843dataset emolialang enspeaker EN_bd3lpFry3JItotal 20.1schain gain +3.7 dBseam step 0.8 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, average recording, quiet background, normal-paced, normally alert, fairly steady, moderate pitch range
(slightly relaxed, frequent disfluency, average clarity, casual)A lot of our transformational travel clients, (low mumble) uhm, included in, in part of their, (ahem) uhm, experiences was that you actually would journal, (ahem) uhm,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.1/10; 9.3s, EN.
EN_bd3lpFry3JI_W000449 · in -18.0 dBFS · gain -2.0 dB · emolia-01858
(contemplation·relaxed, some disfluency, somewhat unclear, casual)As part of that, so you'd have an experience, you know, you go for a hike or maybe you'd see, you'd visit a particular, (ahem) uhm,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contemplation; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 6.2/10; 6.7s, EN.
EN_bd3lpFry3JI_W000450 · in -20.0 dBFS · gain +0.0 dB · emolia-01858
(emotional numbness·slightly relaxed, some disfluency, average clarity, casual)You know, not for profit or something in a particular place and then you would sort of process
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as emotional numbness; style: casual, conversational; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.6/10; 4.6s, EN.
EN_bd3lpFry3JI_W000451 · in -18.5 dBFS · gain -1.5 dB · emolia-01858