Two-sided AB2 after speaker cleaning (WavLM -id >= 0.80 on consecutive pairs AND against the first clip), k=4. This is the only source that ships min_cos_consec / min_cos_anchor.
This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_sc-AB2-k4.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words
are a real non-speech sound, printed where it happens.
The full generated caption for any clip is under “full caption & clip
details”. Its perceived-gender and background-noise clauses were re-rendered from
the numeric buckets, because the versions stored in the corpus index had those two
ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
The tags and captions describe the ORIGINAL clips — they are the
corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the
converted audio are the ones printed in each card's own paragraph and chips.
This chain comes from the two-sided rule: it only counts if both emotions move — Longing down and Fear up — by at least 0.25 each.
The chain starts with Fear clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.26.
At the same time Longing goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.47 (lower than 53 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.22, then +0.01, then +0.03 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 36 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.913 before conversion and 0.892 after — it fell by 0.022. Neighbour-to-neighbour the worst pair went 0.916 → 0.866. (The earlier render, with segment 1 left raw, scores 0.635 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.262 in the original and +0.440 after conversion — 168 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Longing, -0.396 became -0.402.
Quality. Mean predicted overall quality across the segments went 2.97 → 3.06 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.913 → 0.892-0.022identity cos neighbours 0.916 → 0.866d_b rescored +0.262 → +0.440d_a rescored -0.396 → -0.402d_a mined -0.396d_b mined 0.262min_cos_consec (site) 0.9377min_cos_anchor (site) 0.9225dataset emolialang enspeaker EN_lVqmFlKUcw4total 35.6schain gain +3.5 dBseam step 1.2 dBcrossfades 100/100/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(steady, formal, authoritative)By late June 1649, the French and some Christian Hurons built St. Marie II on Christian Island, Isle de Saint Joseph.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 7.8s, EN.
EN_lVqmFlKUcw4_W000162 · in -15.3 dBFS · gain -4.7 dB · emolia-02034
(distress, sadness, fear·fairly steady, newsreading, formal)However, facing starvation, lack of supplies, and constant threats of Iroquois attack, the small Saint-Marie II was abandoned in June 1650. The remaining Hurrans and Jesuits departed for Quebec and Ottawa.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as distress, sadness, fear; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 12.4s, EN.
EN_lVqmFlKUcw4_W000163 · in -16.2 dBFS · gain -3.8 dB · emolia-02034
(fear · fairly steady, formal, newsreading)After a series of epidemics, beginning in 1634, some Huron began to mistrust the Jesuits and accused them of being sorcerers casting spells from their books.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear; style: formal, newsreading; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 9.5s, EN.
EN_lVqmFlKUcw4_W000164 · in -15.9 dBFS · gain -4.1 dB · emolia-02034
(fear, distress, sadness·steady, formal, authoritative)As a result of the Iroquois raids and outbreak of disease, many missionaries, traders, and soldiers died.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, distress, sadness; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 6.3s, EN.
EN_lVqmFlKUcw4_W000165 · in -16.3 dBFS · gain -3.7 dB · emolia-02034
This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Infatuation up — by at least 0.25 each.
The chain starts with Infatuation below average — 0.39, lower than 61 % of clips in this corpus — and ends with it strongly present at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.38.
At the same time Emotional Numbness goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.12, then +0.12, then +0.14 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 32 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.787 before conversion and 0.709 after — it fell by 0.078. Neighbour-to-neighbour the worst pair went 0.787 → 0.725. (The earlier render, with segment 1 left raw, scores 0.579 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.382 in the original and +0.049 after conversion — 13 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Emotional Numbness, -0.382 became -0.306.
Quality. Mean predicted overall quality across the segments went 2.95 → 3.08 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.787 → 0.709-0.078identity cos neighbours 0.787 → 0.725d_b rescored +0.382 → +0.049d_a rescored -0.382 → -0.306d_a mined -0.381d_b mined 0.382min_cos_consec (site) 0.9224min_cos_anchor (site) 0.9180dataset emolialang enspeaker EN_B00067_S00734total 30.5schain gain +2.2 dBseam step 1.6 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, no disfluency
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, fairly narrow pitch, no audible breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, confusion; style: casual, authoritative; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.8/10; 5.1s, EN.
EN_B00067_S00734_W000005 · in -16.5 dBFS · gain -3.5 dB · emolia-01528
(concentration·normal-paced, steady, moderate pitch range, newsreading)Except for the western part of Europe, this type of climate is confined to narrow ranges of occurrences mainly in the low latitudes and to the east of the continents where it appears in the form of
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 10.2s, EN.
EN_B00067_S00734_W000006 · in -17.0 dBFS · gain -3.0 dB · emolia-01528
(normal-paced, fairly steady, moderate pitch range, newsreading)The oceanic climate exists in an arc spreading across the northwestern coast of North America from the Alaskan Panhandle to northern California, in general the coastal areas of the Pacific Northwest.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.9s, EN.
EN_B00067_S00734_W000007 · in -14.1 dBFS · gain -6.0 dB · emolia-01528
(normal-paced, fairly steady, moderate pitch range, formal)The Tristan da Cunha Archipelago in the South Atlantic also has an oceanic climate
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.7/10; 4.9s, EN.
EN_B00067_S00734_W000008 · in -14.1 dBFS · gain -5.9 dB · emolia-01528
This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Fatigue Exhaustion up — by at least 0.25 each.
The chain starts with Fatigue Exhaustion below average — 0.31, lower than 69 % of clips in this corpus — and ends with it clearly present at 0.71, higher than 71 % of clips in this corpus. That is a total rise of 0.40.
At the same time Emotional Numbness goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.58 (higher than 58 % of clips in this corpus), a change of -0.42. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are -0.04, then +0.22, then +0.22 — not a clean run: step 1 moves back the other way by 0.04 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.92 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 32 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.869 before conversion and 0.818 after — it fell by 0.051. Neighbour-to-neighbour the worst pair went 0.852 → 0.770. (The earlier render, with segment 1 left raw, scores 0.713 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.398 in the original and +0.339 after conversion — 85 % of the delta retained, which is most of it. On the other named axis, Emotional Numbness, -0.417 became -0.865.
Quality. Mean predicted overall quality across the segments went 2.94 → 3.00 (+0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.869 → 0.818-0.051identity cos neighbours 0.852 → 0.770d_b rescored +0.398 → +0.339d_a rescored -0.417 → -0.865d_a mined -0.417d_b mined 0.398min_cos_consec (site) 0.9212min_cos_anchor (site) 0.8740dataset emolialang enspeaker EN_j6meOXR6XnEtotal 30.8schain gain +2.2 dBseam step 0.8 dBcrossfades 150/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(emotional numbness · measured, steady, fairly narrow pitch, formal)They are beryllium-B, magnesium-Mg, calcium-Ca, strontium-senia, barium-Ba, and radium-Ra
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, authoritative; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.9/10; 7.6s, EN.
EN_j6meOXR6XnE_W000001 · in -17.3 dBFS · gain -2.7 dB · emolia-00768
(emotional numbness ·normal-paced, fairly steady, moderate pitch range, formal)The elements have very similar properties, they are all shiny, silvery-white, somewhat reactive metals at standard temperature and pressure, structurally, they have in common an outer s-electron shell which is full.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.6s, EN.
EN_j6meOXR6XnE_W000002 · in -18.6 dBFS · gain -1.4 dB · emolia-00768
(normal-paced, fairly steady, moderate pitch range, formal)Experiments have been conducted to attempt the synthesis of element 120, the next potential
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 1.2/10; 7.9s, EN.
EN_j6meOXR6XnE_W000004 · in -18.2 dBFS · gain -1.8 dB · emolia-00768
(normal-paced, fairly steady, moderate pitch range, formal)Most of the chemistry has been observed only for the first five members of the group
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.8/10; 4.3s, EN.
EN_j6meOXR6XnE_W000008 · in -16.9 dBFS · gain -3.1 dB · emolia-00768
This chain comes from the two-sided rule: it only counts if both emotions move — Infatuation down and Doubt up — by at least 0.25 each.
The chain starts with Doubt around average — 0.49, lower than 51 % of clips in this corpus — and ends with it clearly present at 0.75, higher than 75 % of clips in this corpus. That is a total rise of 0.26.
At the same time Infatuation goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.57 (higher than 57 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are -0.12, then +0.19, then +0.19 — not a clean run: step 1 moves back the other way by 0.12 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 19 s · zh · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.876 before conversion and 0.856 after — it fell by 0.020. Neighbour-to-neighbour the worst pair went 0.866 → 0.874. (The earlier render, with segment 1 left raw, scores 0.767 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.263 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Infatuation, -0.322 became +0.014.
Quality. Mean predicted overall quality across the segments went 2.78 → 2.95 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.876 → 0.856-0.020identity cos neighbours 0.866 → 0.874d_b rescored +0.263 → +0.000d_a rescored -0.322 → +0.014d_a mined -0.322d_b mined 0.262min_cos_consec (site) 0.8810min_cos_anchor (site) 0.8757dataset emolialang zhspeaker ZH_B00049_S08693total 17.5schain gain +2.7 dBseam step 0.8 dBcrossfades 100/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a child feminine voice · average recording, measured, normally alert, slightly relaxed, moderate pitch range, light breath
(fairly steady, no disfluency, clear, storytelling)听下面的录音回答。第二小题。
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: storytelling, cartoonish; average recording, no background noise; genuineness 2.2/6; vocal-burst blend 1.7/10; 4.0s, ZH.
ZH_B00049_S08693_W000001 · in -22.0 dBFS · gain +2.0 dB · emolia-03765
(fairly steady, no disfluency, crisply articulate, storytelling)听下面的录音回答。第三小题。
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; crisply articulate, no disfluency, moderate pitch range, light breath; affect is neutral, slightly submissive, slightly guarded; no dominant emotion; style: storytelling, cartoonish; average recording, no background noise; genuineness 2.1/6; vocal-burst blend 1.7/10; 4.2s, ZH.
ZH_B00049_S08693_W000002 · in -22.2 dBFS · gain +2.2 dB · emolia-03765
(fairly steady, no disfluency, average clarity, storytelling)听下面的录音回答。第四小题。
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: storytelling, formal; average recording, no background noise; genuineness 2.3/6; vocal-burst blend 1.7/10; 3.9s, ZH.
ZH_B00049_S08693_W000003 · in -22.6 dBFS · gain +2.6 dB · emolia-03765
(moderately variable, frequent disfluency, average clarity, cartoonish)听到小三个选,每段听后的放两遍,你有想到小题的。
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is slightly cool, slightly dark, slightly rough, thin; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: cartoonish, storytelling; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.6/10; 6.0s, ZH.
ZH_B00049_S08693_W000004 · in -20.9 dBFS · gain +0.9 dB · emolia-03765
This chain comes from the two-sided rule: it only counts if both emotions move — Impatience and Irritability down and Triumph up — by at least 0.25 each.
The chain starts with Triumph around average — 0.56, higher than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.42.
At the same time Impatience and Irritability goes the other way, from 0.62 (higher than 62 % of clips in this corpus) to 0.28 (lower than 72 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.18, then +0.01 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 59 s · fr · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.593 before conversion and 0.618 after — it rose by 0.025. Neighbour-to-neighbour the worst pair went 0.724 → 0.707. (The earlier render, with segment 1 left raw, scores 0.500 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Triumph moved +0.424 in the original and +0.472 after conversion — 111 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Impatience and Irritability, -0.339 became -0.242.
Quality. Mean predicted overall quality across the segments went 2.94 → 3.13 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.593 → 0.618+0.025identity cos neighbours 0.724 → 0.707d_b rescored +0.424 → +0.472d_a rescored -0.339 → -0.242d_a mined -0.339d_b mined 0.424min_cos_consec (site) 0.8511min_cos_anchor (site) 0.8195dataset emolialang frspeaker FR_50L2Zil1Wactotal 57.9schain gain +0.8 dBseam step 1.7 dBcrossfades 150/100/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, slightly dark, balanced body, average recording, quiet background, measured, slightly relaxed, fairly steady
(normally alert, moderate pitch range, light breath, monologue)Alors, on utilise des indicateurs pour différentes unités de temps.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.5/10; 4.3s, FR.
FR_50L2Zil1Wac_W000004 · in -14.3 dBFS · gain -5.7 dB · emolia-02734
(pride, contentment, fatigue exhaustion· normally alert, moderate pitch range, audible breath, monologue)Moi, je regarde les unités de temps à partir du 4 heures jusqu'au 1 minute et même moins puisque maintenant je trade également antique. Donc c'est très, très rapide. Mais les indicateurs que j'utilise vous donnent vraiment des informations très, très pertinentes.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as pride, contentment, fatigue exhaustion; style: monologue, didactic; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 2.5/10; 15.9s, FR.
FR_50L2Zil1Wac_W000005 · in -15.6 dBFS · gain -4.4 dB · emolia-02734
(triumph, pride, concentration·subdued, fairly narrow pitch, light breath, monologue)Et j'utilise cinq indicateurs. Pourquoi? Parce que j'ai fait une sélection de différents indicateurs qui vont me donner une confirmation de ce que dit un, deux ou trois indicateurs. Donc, si j'ai une, (ahem) euh, une convergence de signaux qui me disent, oui, là, le marché, il va probablement monter.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, pride, concentration; style: monologue, didactic; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 1.8/10; 19.4s, FR.
FR_50L2Zil1Wac_W000006 · in -14.9 dBFS · gain -5.1 dB · emolia-02734
(triumph, concentration, pride ·normally alert, moderate pitch range, light breath, monologue)on est jamais sûr à 100%, mais si j'ai une indicateur de 2, 3, 4 indicateurs sur 5 et un indicateur qui va m'indiquer le meilleur moment pour rentrer en position, eh bien, à ce moment-là, effectivement, je rentre en position et j'ai des indicateurs qui vont me dire également quand il est
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, concentration, pride; style: monologue, casual; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 7.2/10; 18.9s, FR.
FR_50L2Zil1Wac_W000007 · in -15.2 dBFS · gain -4.8 dB · emolia-02734
This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Emotional Numbness up — by at least 0.25 each.
The chain starts with Emotional Numbness around average — 0.51, higher than 51 % of clips in this corpus — and ends with it strongly present at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.26.
At the same time Concentration goes the other way, from 0.81 (higher than 81 % of clips in this corpus) to 0.51 (right about the corpus median), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.05, then +0.14, then +0.08 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 35 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.684 before conversion and 0.670 after — it fell by 0.014. Neighbour-to-neighbour the worst pair went 0.684 → 0.670. (The earlier render, with segment 1 left raw, scores 0.598 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.261 in the original and +0.477 after conversion — 183 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.303 became -0.263.
Quality. Mean predicted overall quality across the segments went 2.87 → 3.03 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.684 → 0.670-0.014identity cos neighbours 0.684 → 0.670d_b rescored +0.261 → +0.477d_a rescored -0.303 → -0.263d_a mined -0.302d_b mined 0.261min_cos_consec (site) 0.8101min_cos_anchor (site) 0.8248dataset emolialang enspeaker EN_B00001_S07631total 33.8schain gain +1.3 dBseam step 2.1 dBcrossfades 150/150/100 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, balanced body, slightly relaxed
(measured, normally alert, fairly steady, monologue)So if you want to look at decarbonization, what do we have to do?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 0.2/10; 5.3s, EN.
EN_B00001_S07631_W000052 · in -21.6 dBFS · gain +1.6 dB · emolia-00277
(concentration· measured, very low-energy, fairly steady, monologue)Well, in the United States, the biggest sim- single source of emissions is electricity generation and transportation. Those are the two biggies. So if it's electricity, you've got to talk about fuel switching.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, didactic; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.0/10; 14.7s, EN.
EN_B00001_S07631_W000053 · in -24.4 dBFS · gain +4.4 dB · emolia-00277
(slow, very low-energy, steady, whispered)So we talk about renewables like solar and wind and nuclear as well as things like geothermal and, (low mumble) uh,
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: whispered, monologue; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 0.1/10; 10.8s, EN.
EN_B00001_S07631_W000054 · in -28.6 dBFS · gain +8.6 dB · emolia-00277
(measured, normally alert, steady, monologue)Estimate on the typical driving cycle is 12%.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 0.1/10; 3.5s, EN.
EN_B00001_S07631_W000055 · in -29.4 dBFS · gain +9.4 dB · emolia-00277
This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Contemplation up — by at least 0.25 each.
The chain starts with Contemplation clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.28.
At the same time Emotional Numbness goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are -0.05, then +0.09, then +0.24 — not a clean run: step 1 moves back the other way by 0.05 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.90 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 36 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.851 before conversion and 0.793 after — it fell by 0.057. Neighbour-to-neighbour the worst pair went 0.715 → 0.669. (The earlier render, with segment 1 left raw, scores 0.711 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.275 in the original and +0.277 after conversion — 101 % of the delta retained, which is essentially all of it. On the other named axis, Emotional Numbness, -0.258 became -0.184.
Quality. Mean predicted overall quality across the segments went 2.89 → 3.05 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.851 → 0.793-0.057identity cos neighbours 0.715 → 0.669d_b rescored +0.275 → +0.277d_a rescored -0.258 → -0.184d_a mined -0.258d_b mined 0.275min_cos_consec (site) 0.8976min_cos_anchor (site) 0.8747dataset emolialang enspeaker EN_B00039_S03967total 34.7schain gain +2.5 dBseam step 1.0 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(emotional numbness · fairly steady, little disfluency, average clarity, monologue)That those effective microorganisms end up adding to the soils and so literally it's possible to combine
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness; style: monologue, authoritative; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 1.4/10; 6.6s, EN.
EN_B00039_S03967_W000169 · in -22.7 dBFS · gain +2.7 dB · emolia-01005
(steady, almost no disfluency, clear, formal)These effective amendments, these soil amendments, to restore depleted soils back to productivity and do what we call extensive agriculture.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.5/10; 7.0s, EN.
EN_B00039_S03967_W000170 · in -24.2 dBFS · gain +4.2 dB · emolia-01005
(fairly steady, little disfluency, average clarity, monologue)Now there's several definitions for extensive agriculture, but for me it's about
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 1.4/10; 3.5s, EN.
EN_B00039_S03967_W000171 · in -22.0 dBFS · gain +2.0 dB · emolia-01005
(contemplation, concentration, relief·steady, some disfluency, average clarity, monologue)Taking land that was used in the past and is no longer arable and restoring it to productivity. This is an opportunity for us to really make a difference in agriculture and insuring food security on earth. And at the same time, all that biochar that we're putting into the soil is staying for centuries.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, concentration, relief; style: monologue, formal; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.0/10; 18.1s, EN.
EN_B00039_S03967_W000172 · in -25.2 dBFS · gain +5.2 dB · emolia-01005
Intoxication Altered States of Consciousness ↓ / Shame ↑identity −0.02emotion 90 % sc-AB2-k4 · #8
This chain comes from the two-sided rule: it only counts if both emotions move — Intoxication Altered States of Consciousness down and Shame up — by at least 0.25 each.
The chain starts with Shame barely there — 0.25, lower than 75 % of clips in this corpus — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.63.
At the same time Intoxication Altered States of Consciousness goes the other way, from 0.65 (higher than 65 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.19, then +0.22, then +0.22 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.91 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 26 s · zh · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.764 before conversion and 0.743 after — it fell by 0.022. Neighbour-to-neighbour the worst pair went 0.874 → 0.857. (The earlier render, with segment 1 left raw, scores 0.688 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.633 in the original and +0.570 after conversion — 90 % of the delta retained, which is essentially all of it. On the other named axis, Intoxication Altered States of Consciousness, -0.286 became -0.397.
Quality. Mean predicted overall quality across the segments went 3.02 → 3.12 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.764 → 0.743-0.022identity cos neighbours 0.874 → 0.857d_b rescored +0.633 → +0.570d_a rescored -0.286 → -0.397d_a mined -0.286d_b mined 0.633min_cos_consec (site) 0.9051min_cos_anchor (site) 0.8316dataset emolialang zhspeaker ZH_B00043_S09998total 24.5schain gain +0.8 dBseam step 0.5 dBcrossfades 150/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a child feminine voice · neutral-toned, neutral-bright, fairly smooth, no background noise, normally alert, slightly relaxed, no disfluency, clear
This chain comes from the two-sided rule: it only counts if both emotions move — Thankfulness Gratitude down and Concentration up — by at least 0.25 each.
The chain starts with Concentration clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.36.
At the same time Thankfulness Gratitude goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.62 (higher than 62 % of clips in this corpus), a change of -0.35. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.19, then +0.09, then +0.08 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 45 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.796 before conversion and 0.777 after — it fell by 0.018. Neighbour-to-neighbour the worst pair went 0.886 → 0.875. (The earlier render, with segment 1 left raw, scores 0.578 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.361 in the original and +0.312 after conversion — 86 % of the delta retained, which is most of it. On the other named axis, Thankfulness Gratitude, -0.355 became -0.265.
Quality. Mean predicted overall quality across the segments went 2.45 → 3.02 (+0.56) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.796 → 0.777-0.018identity cos neighbours 0.886 → 0.875d_b rescored +0.361 → +0.312d_a rescored -0.355 → -0.265d_a mined -0.354d_b mined 0.363min_cos_consec (site) 0.9011min_cos_anchor (site) 0.9011dataset emolialang enspeaker EN_K73s99ixjGutotal 44.3schain gain +4.5 dBseam step 1.5 dBcrossfades 150/100/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, average recording, normally alert, fairly steady, moderate pitch range
(thankfulness gratitude, pain · normal-paced, slightly relaxed, little disfluency, monologue)In our county, in Monterey County, (ahem) uhm, we're under a county organized health system. So we're the sole health plan.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as thankfulness gratitude, pain; style: monologue, formal; average recording, no background noise; genuineness 1.9/6; vocal-burst blend 1.0/10; 5.8s, EN.
EN_K73s99ixjGu_W000305 · in -19.9 dBFS · gain -0.1 dB · emolia-02493
(thankfulness gratitude · normal-paced, slightly relaxed, some disfluency, formal)And has been referenced, we're a non-profit health plan, so our focus as I mentioned before is ensuring that our members have access to those quality services and that we're partnering with our local providers.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as thankfulness gratitude; style: formal; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 1.5/10; 10.8s, EN.
EN_K73s99ixjGu_W000307 · in -21.2 dBFS · gain +1.2 dB · emolia-02493
(fear, concentration·brisk, neutral tension, some disfluency, formal)And so I'm sitting on a panel with hospitals that have partnered with our health plan to implement, (ahem) um, EV navigators in our hospitals. And so in 2018, our health plan will partner with the hospitals to have navigators available in the emergency department to help our members access the system. So it's not just about the medical care, but are there other services that they would need?
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, minimal breath; affect is mildly positive, neutral stance, slightly guarded; reads as fear, concentration; style: formal, monologue; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 2.2/10; 17.9s, EN.
EN_K73s99ixjGu_W000309 · in -21.4 dBFS · gain +1.4 dB · emolia-02493
(concentration, doubt, interest·normal-paced, slightly relaxed, little disfluency, formal)How do we make the healthcare delivery system better? How do we ensure quality and be responsive to what people are requesting? So I wanted to point that out about the MediCal system and our provider network (ahem) generally here.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as concentration, doubt, interest; style: formal, monologue; average recording, no background noise; genuineness 2.3/6; vocal-burst blend 2.5/10; 10.5s, EN.
EN_K73s99ixjGu_W000311 · in -23.1 dBFS · gain +3.1 dB · emolia-02493
This chain comes from the two-sided rule: it only counts if both emotions move — Jealousy and Envy down and Contemplation up — by at least 0.25 each.
The chain starts with Contemplation clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.29.
At the same time Jealousy and Envy goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.59 (higher than 59 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.13, then -0.08, then +0.24 — not a clean run: step 2 moves back the other way by 0.08 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.91 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 58 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.885 before conversion and 0.834 after — it fell by 0.050. Neighbour-to-neighbour the worst pair went 0.902 → 0.834. (The earlier render, with segment 1 left raw, scores 0.680 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.289 in the original and +0.268 after conversion — 93 % of the delta retained, which is essentially all of it. On the other named axis, Jealousy and Envy, -0.397 became -0.339.
Quality. Mean predicted overall quality across the segments went 2.88 → 3.13 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.885 → 0.834-0.050identity cos neighbours 0.902 → 0.834d_b rescored +0.289 → +0.268d_a rescored -0.397 → -0.339d_a mined -0.398d_b mined 0.289min_cos_consec (site) 0.9052min_cos_anchor (site) 0.8724dataset emolialang enspeaker EN_yOxdZFRyA3Ytotal 57.1schain gain +6.0 dBseam step 0.2 dBcrossfades 150/100/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, slightly bright, fairly smooth, balanced body, good recording, normally alert, slightly relaxed, some disfluency
(jealousy and envy, emotional numbness, longing · normal-paced, fairly steady, moderate pitch range, casual)And now lives in his big manor house far away from Mexico City where Noemi herself lives. And Noemi is the one who has to figure out what's going on. And she is in fact very strong.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as jealousy and envy, emotional numbness, longing; style: casual, conversational; good recording, quiet background; genuineness 3.4/6; vocal-burst blend 4.2/10; 11.2s, EN.
EN_yOxdZFRyA3Y_W000060 · in -23.7 dBFS · gain +3.7 dB · emolia-02624
(sourness· normal-paced, fairly steady, moderate pitch range, casual)In terms of her personality, and she also is a (ahem) well-developed character with a lot of different characteristics as opposed to just being a stereotype or a cardboard cutout or existing to serve some other character's storyline.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as sourness; style: casual, monologue; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 0.0/10; 16.1s, EN.
EN_yOxdZFRyA3Y_W000061 · in -23.6 dBFS · gain +3.6 dB · emolia-02624
(interest, contentment, concentration·brisk, moderately variable, wide pitch range, casual)Writing style is another element, and this can refer to the complexity of language, the level of detail, and there are some other features that if you check out that novelist guide you will see in there. And Mexican Gothic is compelling, so this is the kind of book where though things are revealed,
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as interest, contentment, concentration; style: casual, dramatic; good recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.7/10; 17.4s, EN.
EN_yOxdZFRyA3Y_W000062 · in -25.2 dBFS · gain +5.2 dB · emolia-02624
(contemplation·normal-paced, fairly steady, moderate pitch range, casual)Sort of deliberately and a little bit at a time, there's always something that if you're the kind of reader this book is going to work for will edge you on to keep reading. So that's an example of one type of writing style.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contemplation; style: casual, monologue; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.5/10; 12.9s, EN.
EN_yOxdZFRyA3Y_W000063 · in -19.3 dBFS · gain -0.7 dB · emolia-02624
This chain comes from the two-sided rule: it only counts if both emotions move — Longing down and Fatigue Exhaustion up — by at least 0.25 each.
The chain starts with Fatigue Exhaustion around average — 0.52, higher than 52 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.33.
At the same time Longing goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.03, then +0.14, then +0.16 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.88 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 31 s · zh · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.729 before conversion and 0.716 after — it fell by 0.014. Neighbour-to-neighbour the worst pair went 0.719 → 0.716. (The earlier render, with segment 1 left raw, scores 0.690 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.328 in the original and +0.253 after conversion — 77 % of the delta retained, which is most of it. On the other named axis, Longing, -0.344 became -0.381.
Quality. Mean predicted overall quality across the segments went 3.07 → 3.13 (+0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.729 → 0.716-0.014identity cos neighbours 0.719 → 0.716d_b rescored +0.328 → +0.253d_a rescored -0.344 → -0.381d_a mined -0.344d_b mined 0.328min_cos_consec (site) 0.8818min_cos_anchor (site) 0.9016dataset emolialang zhspeaker ZH_B00062_S04966total 30.0schain gain -2.0 dBseam step 1.1 dBcrossfades 100/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, no disfluency
This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Doubt up — by at least 0.25 each.
The chain starts with Doubt clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.27.
At the same time Concentration goes the other way, from 0.80 (higher than 80 % of clips in this corpus) to 0.50 (right about the corpus median), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.17, then -0.06, then +0.16 — not a clean run: step 2 moves back the other way by 0.06 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 30 s · zh · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.869 before conversion and 0.913 after — it rose by 0.044. Neighbour-to-neighbour the worst pair went 0.869 → 0.913. (The earlier render, with segment 1 left raw, scores 0.818 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.273 in the original and +0.329 after conversion — 121 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.301 became -0.208.
Quality. Mean predicted overall quality across the segments went 3.09 → 3.23 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.869 → 0.913+0.044identity cos neighbours 0.869 → 0.913d_b rescored +0.273 → +0.329d_a rescored -0.301 → -0.208d_a mined -0.302d_b mined 0.273min_cos_consec (site) 0.9036min_cos_anchor (site) 0.9272dataset emolialang zhspeaker ZH_B00016_S02373total 29.4schain gain +0.8 dBseam step 0.4 dBcrossfades 100/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady
This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Fatigue Exhaustion up — by at least 0.25 each.
The chain starts with Fatigue Exhaustion below average — 0.31, lower than 69 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.52.
At the same time Concentration goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.21, then +0.19, then +0.13 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.92 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 46 s · zh · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.805 before conversion and 0.883 after — it rose by 0.078. Neighbour-to-neighbour the worst pair went 0.910 → 0.859. (The earlier render, with segment 1 left raw, scores 0.784 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.524 in the original and +0.466 after conversion — 89 % of the delta retained, which is most of it. On the other named axis, Concentration, -0.318 became -0.366.
Quality. Mean predicted overall quality across the segments went 3.00 → 3.31 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.805 → 0.883+0.078identity cos neighbours 0.910 → 0.859d_b rescored +0.524 → +0.466d_a rescored -0.318 → -0.366d_a mined -0.319d_b mined 0.524min_cos_consec (site) 0.9155min_cos_anchor (site) 0.8529dataset emolialang zhspeaker ZH_B00059_S07231total 44.9schain gain +0.8 dBseam step 1.5 dBcrossfades 100/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, normally alert, slightly relaxed, fairly steady
This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Contemplation up — by at least 0.25 each.
The chain starts with Contemplation clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.26.
At the same time Concentration goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.18, then -0.16, then +0.24 — not a clean run: step 2 moves back the other way by 0.16 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 41 s · zh · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.848 before conversion and 0.857 after — it rose by 0.009. Neighbour-to-neighbour the worst pair went 0.848 → 0.857. (The earlier render, with segment 1 left raw, scores 0.755 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.257 in the original and +0.098 after conversion — 38 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Concentration, -0.271 became -0.448.
Quality. Mean predicted overall quality across the segments went 3.03 → 3.28 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.848 → 0.857+0.009identity cos neighbours 0.848 → 0.857d_b rescored +0.257 → +0.098d_a rescored -0.271 → -0.448d_a mined -0.270d_b mined 0.256min_cos_consec (site) 0.8660min_cos_anchor (site) 0.8660dataset emolialang zhspeaker ZH_B00006_S08924total 40.1schain gain -0.5 dBseam step 1.3 dBcrossfades 100/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, no background noise, measured, normally alert, slightly relaxed
(concentration, pride · steady, clear, fairly narrow pitch, didactic)Dac技术也可以与CCS技术结合使用,对CCS技术储存中泄露的二氧化碳进行普及,从而提进一步提高碳普及的能力。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, pride; style: didactic, monologue; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 0.2/10; 13.2s, ZH.
ZH_B00006_S08924_W000090 · in -19.6 dBFS · gain -0.4 dB · emolia-03336
This chain comes from the two-sided rule: it only counts if both emotions move — Pain down and Emotional Numbness up — by at least 0.25 each.
The chain starts with Emotional Numbness clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.29.
At the same time Pain goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.35 (lower than 65 % of clips in this corpus), a change of -0.37. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.12, then +0.15, then +0.01 — a plateau around step 3, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 32 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.941 before conversion and 0.852 after — it fell by 0.089. Neighbour-to-neighbour the worst pair went 0.934 → 0.852. (The earlier render, with segment 1 left raw, scores 0.535 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.285 in the original and +0.313 after conversion — 110 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Pain, -0.369 became -0.369.
Quality. Mean predicted overall quality across the segments went 2.96 → 3.05 (+0.09) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.941 → 0.852-0.089identity cos neighbours 0.934 → 0.852d_b rescored +0.285 → +0.313d_a rescored -0.369 → -0.369d_a mined -0.369d_b mined 0.285min_cos_consec (site) 0.9635min_cos_anchor (site) 0.9295dataset emolialang enspeaker EN_lU9zMRQqiN8total 30.7schain gain +1.5 dBseam step 0.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fairly steady, formal, authoritative)Provisions for digital signatures in accordance with the PDF Advanced Electronic Signatures – PADES standard
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.2/10; 6.2s, EN.
EN_lU9zMRQqiN8_W000042 · in -15.3 dBFS · gain -4.7 dB · emolia-01484
(fairly steady, formal, newsreading)The option of embedding PDF.of files to facilitate archiving of sets of documents with a single file, Part 2 defines three conformance levels
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.2/10; 8.5s, EN.
EN_lU9zMRQqiN8_W000043 · in -15.2 dBFS · gain -4.8 dB · emolia-01484
(emotional numbness, intoxication altered states of consciousness·steady, formal, authoritative)PDF, A2A and PDF, A2B correspond to conformance levels A and B in PDF, A1
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, intoxication altered states of consciousness; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.0/10; 6.6s, EN.
EN_lU9zMRQqiN8_W000044 · in -14.3 dBFS · gain -5.7 dB · emolia-01484
(emotional numbness ·fairly steady, formal, authoritative)A new conformance level, pdf.a2u, represents level B conformance, pdf.a2b, with the additional requirement that all text in the document have Unicode mapping
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.1s, EN.
EN_lU9zMRQqiN8_W000045 · in -15.2 dBFS · gain -4.8 dB · emolia-01484
Concentration ↓ / Intoxication Altered States of Consciousness ↑identity +0.00emotion 121 % sc-AB2-k4 · #16
This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Intoxication Altered States of Consciousness up — by at least 0.25 each.
The chain starts with Intoxication Altered States of Consciousness around average — 0.51, right about the corpus median — and ends with it strongly present at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.29.
At the same time Concentration goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.43. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.02, then +0.09, then +0.18 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 28 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.764 before conversion and 0.767 after — it rose by 0.002. Neighbour-to-neighbour the worst pair went 0.806 → 0.756. (The earlier render, with segment 1 left raw, scores 0.637 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.289 in the original and +0.349 after conversion — 121 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.429 became -0.459.
Quality. Mean predicted overall quality across the segments went 2.65 → 2.83 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.764 → 0.767+0.002identity cos neighbours 0.806 → 0.756d_b rescored +0.289 → +0.349d_a rescored -0.429 → -0.459d_a mined -0.430d_b mined 0.289min_cos_consec (site) 0.8454min_cos_anchor (site) 0.8201dataset emolialang enspeaker EN_TWtwMbCa0zMtotal 26.8schain gain +2.8 dBseam step 0.9 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-bright, balanced body, quiet background, normally alert, slightly relaxed, fairly steady, average clarity, moderate pitch range
(concentration · brisk, some disfluency, casual, authoritative)Otherwise the bond will be keep on breaking. This is called osteoporosis. We have injectable bond cement for those kind of system. Inject it so that the bond will be filled.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration; style: casual, authoritative; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 2.1/10; 9.2s, EN.
EN_TWtwMbCa0zM_W000339 · in -25.2 dBFS · gain +5.2 dB · emolia-02241
(brisk, some disfluency, authoritative, monologue)Then those, (ahem) these are the nano fiber, the bond screws, like I mentioned before. Use this material for as a normal screw for your bond breakage, then it will dissolve back.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, monologue; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.6/10; 9.1s, EN.
EN_TWtwMbCa0zM_W000340 · in -23.6 dBFS · gain +3.6 dB · emolia-02241
(brisk, some disfluency, casual, authoritative)So there are lots and lots of varieties of systems working. So first we will go through the bonds or joints and teeth.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, authoritative; below-average recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.7/10; 5.5s, EN.
EN_TWtwMbCa0zM_W000341 · in -21.9 dBFS · gain +1.9 dB · emolia-02241
(normal-paced, frequent disfluency, casual, authoritative)We have a lot of options to make the bonds from.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, authoritative; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 1.1/10; 3.6s, EN.
EN_TWtwMbCa0zM_W000342 · in -22.9 dBFS · gain +2.9 dB · emolia-02241
This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Shame up — by at least 0.25 each.
The chain starts with Shame below average — 0.41, lower than 59 % of clips in this corpus — and ends with it strongly present at 0.79, higher than 79 % of clips in this corpus. That is a total rise of 0.38.
At the same time Contemplation goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.39 (lower than 61 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.06, then +0.19, then +0.12 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 22 s · zh · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.858 before conversion and 0.883 after — it rose by 0.025. Neighbour-to-neighbour the worst pair went 0.858 → 0.883. (The earlier render, with segment 1 left raw, scores 0.826 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.378 in the original and +0.230 after conversion — 61 % of the delta retained. On the other named axis, Contemplation, -0.336 became -0.153.
Quality. Mean predicted overall quality across the segments went 2.94 → 3.04 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.858 → 0.883+0.025identity cos neighbours 0.858 → 0.883d_b rescored +0.378 → +0.230d_a rescored -0.336 → -0.153d_a mined -0.336d_b mined 0.378min_cos_consec (site) 0.8670min_cos_anchor (site) 0.8670dataset emolialang zhspeaker ZH_B00044_S05913total 21.4schain gain +2.5 dBseam step 1.8 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, measured, normally alert
(narration, authoritative)王明作如何继续全国抗战和争取抗战胜利呢的报告。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, authoritative; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 4.6/10; 4.9s, ZH.
ZH_B00044_S05913_W000007 · in -26.2 dBFS · gain +6.2 dB · emolia-03715
This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Fatigue Exhaustion up — by at least 0.25 each.
The chain starts with Fatigue Exhaustion below average — 0.33, lower than 67 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.52.
At the same time Concentration goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.09, then +0.25, then +0.19 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 40 s · zh · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.871 before conversion and 0.887 after — it rose by 0.017. Neighbour-to-neighbour the worst pair went 0.889 → 0.916. (The earlier render, with segment 1 left raw, scores 0.819 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.525 in the original and +0.700 after conversion — 133 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.296 became -0.306.
Quality. Mean predicted overall quality across the segments went 3.01 → 3.24 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.871 → 0.887+0.017identity cos neighbours 0.889 → 0.916d_b rescored +0.525 → +0.700d_a rescored -0.296 → -0.306d_a mined -0.297d_b mined 0.525min_cos_consec (site) 0.9042min_cos_anchor (site) 0.9199dataset emolialang zhspeaker ZH_B00028_S06080total 39.5schain gain +2.8 dBseam step 1.4 dBcrossfades 100/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · fairly smooth, balanced body, no background noise, measured, normally alert, slightly relaxed, fairly steady, no disfluency
(concentration, pride · fairly narrow pitch, monologue, narration)在距离秦王嬴政统一六国一百年前,楚国的疆域北设黄河,东到江浙,西控巴蜀,南至芈月,是战国七雄中领土最大的国家,军事力量也是称雄一世。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, pride; style: monologue, narration; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 5.3/10; 14.7s, ZH.
ZH_B00028_S06080_W000001 · in -22.8 dBFS · gain +2.8 dB · emolia-03552
This chain comes from the two-sided rule: it only counts if both emotions move — Astonishment Surprise down and Impatience and Irritability up — by at least 0.25 each.
The chain starts with Impatience and Irritability clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.29.
At the same time Astonishment Surprise goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.09, then +0.15, then +0.05 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 28 s · zh · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.822 before conversion and 0.819 after — it fell by 0.002. Neighbour-to-neighbour the worst pair went 0.795 → 0.782. (The earlier render, with segment 1 left raw, scores 0.755 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.286 in the original and +0.600 after conversion — 210 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Astonishment Surprise, -0.275 became -0.238.
Quality. Mean predicted overall quality across the segments went 2.89 → 3.10 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.822 → 0.819-0.002identity cos neighbours 0.795 → 0.782d_b rescored +0.286 → +0.600d_a rescored -0.275 → -0.238d_a mined -0.274d_b mined 0.286min_cos_consec (site) 0.8112min_cos_anchor (site) 0.8199dataset emolialang zhspeaker ZH_B00060_S09119total 26.6schain gain +2.3 dBseam step 0.5 dBcrossfades 150/100/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, fast, moderately variable, some disfluency
This chain comes from the two-sided rule: it only counts if both emotions move — Pride down and Contemplation up — by at least 0.25 each.
The chain starts with Contemplation clearly present — 0.59, higher than 59 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.37.
At the same time Pride goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are -0.02, then +0.25, then +0.14 — not a clean run: step 1 moves back the other way by 0.02 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 30 s · fr · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.726 before conversion and 0.690 after — it fell by 0.036. Neighbour-to-neighbour the worst pair went 0.796 → 0.693. (The earlier render, with segment 1 left raw, scores 0.602 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.370 in the original and +0.319 after conversion — 86 % of the delta retained, which is most of it. On the other named axis, Pride, -0.378 became -0.490.
Quality. Mean predicted overall quality across the segments went 2.88 → 3.08 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.726 → 0.690-0.036identity cos neighbours 0.796 → 0.693d_b rescored +0.370 → +0.319d_a rescored -0.378 → -0.490d_a mined -0.378d_b mined 0.371min_cos_consec (site) 0.8394min_cos_anchor (site) 0.8384dataset emolialang frspeaker FR_Z1mCTAgIQ6Itotal 28.9schain gain +1.0 dBseam step 1.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range, light breath
(pride · measured, frequent disfluency, somewhat unclear, monologue)Donc, ce programme, maintenant, on l'entend vraiment, (low mumble) euh, depuis un bon nombre d'années dans toutes les disciplines académiques, (low mumble) euh, et il raisonne ce programme des, des Amériques à l'Europe.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride; style: monologue, didactic; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 0.0/10; 10.2s, FR.
FR_Z1mCTAgIQ6I_W000035 · in -20.6 dBFS · gain +0.6 dB · emolia-02652
(normal-paced, some disfluency, average clarity, conversational)Alors, comme je vais le montrer, en fait, ce, ce, ce programme qui fait fond sur un diagnostic a d'abord été articulé par des philosophes américains, même si maintenant on le retrouve partout.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational, monologue; good recording, quiet background; genuineness 2.6/6; vocal-burst blend 1.4/10; 7.9s, FR.
FR_Z1mCTAgIQ6I_W000036 · in -17.3 dBFS · gain -2.7 dB · emolia-02652
(normal-paced, some disfluency, average clarity, formal)beaucoup en anthropologie, et même curieusement en anthropologie des sciences.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 2.7/6; vocal-burst blend 1.5/10; 3.2s, FR.
FR_Z1mCTAgIQ6I_W000037 · in -17.0 dBFS · gain -3.0 dB · emolia-02652
(contemplation·measured, some disfluency, somewhat unclear, monologue)que la modernité aurait privé la religion de son énergie en la réduisant à n'être qu'un ustensile de l'âme. Et donc, c'est même Bruno Latour pouvait dire en 2001
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation; style: monologue, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 1.6/10; 8.2s, FR.
FR_Z1mCTAgIQ6I_W000038 · in -18.3 dBFS · gain -1.7 dB · emolia-02652