proxy_spearman__PXR__T0.60__C0.25__INTERNAL — voice-corrected

Manifest tier. proxy_spearman, rule PXR, T=0.6, step cap 0.25. Population 70 chains (1 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 66.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_proxy_spearman__PXR__T0.60__C0.25__INTERNAL.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
77segments re-voiced
0.672 → 0.685median worst-to-anchor identity cosine
95 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Contemplation ↓  /  Infatuationidentity +0.09 emotion 16 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #1

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation barely there — 0.25, lower than 75 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.62.

At the same time Contemplation goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.30 (lower than 70 % of clips in this corpus), a change of -0.61. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.09, then +0.20, then +0.14 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.51 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.51 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.51, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 27 s · ja · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.534 before conversion and 0.625 after — it rose by 0.092. Neighbour-to-neighbour the worst pair went 0.534 → 0.661. (The earlier render, with segment 1 left raw, scores 0.298 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.623 in the original and +0.097 after conversion — 16 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Contemplation, -0.605 became -0.576.

Quality. Mean predicted overall quality across the segments went 2.82 → 3.15 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.534 → 0.625 +0.092identity cos neighbours 0.534 → 0.661d_b rescored +0.623 → +0.097d_a rescored -0.605 → -0.576d_a mined -0.605d_b mined 0.623min_cos_consec (site) 0.5092min_cos_anchor (site) 0.5092dataset emolialang jaspeaker JA_B00004_S00819total 25.8schain gain -1.4 dBseam step 1.4 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, measured, normally alert, slightly relaxed, moderate pitch range, light breath
(contemplation · steady, little disfluency, average clarity, formal) 日本語入力にはOSを問わず、ITXモズクをおすすめしています。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation; style: formal, authoritative; average recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.6/10; 6.0s, JA.
JA_B00004_S00819_W000046 · in -17.2 dBFS · gain -2.8 dB · emolia-03002
(fairly steady, some disfluency, average clarity, storytelling) この辞書はどう強化されているかわかりませんので試しようがない。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: storytelling, monologue; average recording, no background noise; genuineness 2.0/6; vocal-burst blend 3.9/10; 5.3s, JA.
JA_B00004_S00819_W000047 · in -15.5 dBFS · gain -4.5 dB · emolia-03002
(emotional numbness · fairly steady, little disfluency, clear, monologue) その時代に、携帯素解析による品質判定や意味解析をやってたんです。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: monologue, didactic; average recording, no background noise; genuineness 0.5/6; vocal-burst blend 1.5/10; 6.7s, JA.
JA_B00004_S00819_W000048 · in -17.9 dBFS · gain -2.1 dB · emolia-03002
(confusion · fairly steady, some disfluency, somewhat unclear, casual) 画面のGoogleドライブのファイルをダウンロードします。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion; style: casual, authoritative; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 1.7/10; 4.3s, JA.
JA_B00004_S00819_W000049 · in -13.8 dBFS · gain -6.2 dB · emolia-03002
(steady, no disfluency, clear, authoritative) 次は、Google 日本語入力の強化辞書です。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 2.6/10; 4.3s, JA.
JA_B00004_S00819_W000050 · in -13.7 dBFS · gain -6.3 dB · emolia-03002
Emotional Numbness ↓  /  Interestidentity −0.04 emotion 107 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #2

This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Interest below average — 0.32, lower than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.68.

At the same time Emotional Numbness goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.32 (lower than 68 % of clips in this corpus), a change of -0.61. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then -0.00, then +0.20, then +0.24 — not a clean run: step 2 moves back the other way by 0.00 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.57 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.72 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.57, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 41 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.618 before conversion and 0.580 after — it fell by 0.038. Neighbour-to-neighbour the worst pair went 0.698 → 0.509. (The earlier render, with segment 1 left raw, scores 0.422 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.677 in the original and +0.724 after conversion — 107 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, -0.608 became -0.603.

Quality. Mean predicted overall quality across the segments went 2.49 → 2.99 (+0.50) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.618 → 0.580 -0.038identity cos neighbours 0.698 → 0.509d_b rescored +0.677 → +0.724d_a rescored -0.608 → -0.603d_a mined -0.608d_b mined 0.678min_cos_consec (site) 0.7161min_cos_anchor (site) 0.5688dataset emolialang enspeaker EN_4mNN7vaOCn4total 39.7schain gain +3.5 dBseam step 1.2 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, some disfluency
(emotional numbness · steady, monologue, formal) It's not said that every service in exchange is free to every user.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as emotional numbness; style: monologue, formal; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 0.5/10; 6.0s, EN.
EN_4mNN7vaOCn4_W000116 · in -18.6 dBFS · gain -1.4 dB · emolia-01228
(concentration, doubt · steady, didactic, monologue) Some of the services may be (low mumble) against payment. Some of the services may be restricted to particular subsets of users.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, doubt; style: didactic, monologue; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 0.0/10; 9.7s, EN.
EN_4mNN7vaOCn4_W000117 · in -20.0 dBFS · gain +0.0 dB · emolia-01228
(fairly steady, monologue, casual) But the difference will be is that they will be visible, they will be there, and there can be requests to share them as well.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 0.6/10; 6.3s, EN.
EN_4mNN7vaOCn4_W000118 · in -21.5 dBFS · gain +1.5 dB · emolia-01228
(fairly steady, monologue, casual) Now we are looking at different business models, so how we can essentially
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 1.6/10; 5.7s, EN.
EN_4mNN7vaOCn4_W000119 · in -17.0 dBFS · gain -3.0 dB · emolia-01228
(interest, concentration · fairly steady, monologue, didactic) (low mumble) And we're looking at different ways in which that's happening today. There are examples. If one looks at things like Euro HPC, there is agreements already so that users from other countries
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, concentration; style: monologue, didactic; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 0.3/10; 12.8s, EN.
EN_4mNN7vaOCn4_W000121 · in -18.3 dBFS · gain -1.7 dB · emolia-01228
Disgust ↓  /  Reliefidentity −0.03 emotion 105 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #3

This chain comes from the proxy rule: the same two-sided test as above, but because Relief is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Relief below average — 0.38, lower than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.60.

At the same time Disgust goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.33 (lower than 67 % of clips in this corpus), a change of -0.66. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.14, then +0.23 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 26 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.807 before conversion and 0.782 after — it fell by 0.025. Neighbour-to-neighbour the worst pair went 0.712 → 0.608. (The earlier render, with segment 1 left raw, scores 0.702 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.600 in the original and +0.631 after conversion — 105 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Disgust, -0.664 became -0.414.

Quality. Mean predicted overall quality across the segments went 2.81 → 2.95 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.807 → 0.782 -0.025identity cos neighbours 0.712 → 0.608d_b rescored +0.600 → +0.631d_a rescored -0.664 → -0.414d_a mined -0.664d_b mined 0.600min_cos_consec (site) 0.8557min_cos_anchor (site) 0.8387dataset emolialang enspeaker EN_AVkV0EjES-Ytotal 25.2schain gain +1.0 dBseam step 1.5 dBcrossfades 100/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a child feminine voice · balanced body, no background noise, slightly relaxed, steady, light breath
(disgust · measured, energised, no disfluency, didactic) Hence it is best to avoid junk foods in parties also.
full caption & clip details
A child feminine voice; delivery is energised, measured, slightly relaxed, steady; timbre is slightly cool, slightly bright, fairly smooth, balanced body; very clear, no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as disgust; style: didactic, dramatic; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.6/10; 5.2s, EN.
EN_AVkV0EjES-Y_W000092 · in -17.4 dBFS · gain -2.6 dB · emolia-02626
(emotional numbness · measured, normally alert, no disfluency, formal) Harmful effects of junk food are explained in detail in another tutorial.
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, didactic; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.1/10; 5.9s, EN.
EN_AVkV0EjES-Y_W000093 · in -17.7 dBFS · gain -2.3 dB · emolia-02626
(measured, normally alert, no disfluency, formal) Please visit our website for more details.
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; style: formal, didactic; good recording, no background noise; no dominant emotion; genuineness 0.9/6; vocal-burst blend 0.3/10; 3.0s, EN.
EN_AVkV0EjES-Y_W000094 · in -17.0 dBFS · gain -3.0 dB · emolia-02626
(relief, thankfulness gratitude · slow, normally alert, almost no disfluency, didactic) Food during a celebration or at a party does not have to be unhealthy. With a little effort and planning, healthy food can be served.
full caption & clip details
A child feminine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, thankfulness gratitude; style: didactic, formal; average recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.4/10; 11.6s, EN.
EN_AVkV0EjES-Y_W000095 · in -18.4 dBFS · gain -1.6 dB · emolia-02626
Relief ↓  /  Fatigue Exhaustionidentity +0.04 emotion 102 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #4

This chain comes from the proxy rule: the same two-sided test as above, but because Fatigue Exhaustion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fatigue Exhaustion barely there — 0.15, lower than 85 % of clips in this corpus — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.67.

At the same time Relief goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.24 (lower than 76 % of clips in this corpus), a change of -0.65. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.17, then +0.10, then +0.16 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 46 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.734 before conversion and 0.771 after — it rose by 0.036. Neighbour-to-neighbour the worst pair went 0.734 → 0.771. (The earlier render, with segment 1 left raw, scores 0.730 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.674 in the original and +0.688 after conversion — 102 % of the delta retained, which is essentially all of it. On the other named axis, Relief, -0.645 became -0.466.

Quality. Mean predicted overall quality across the segments went 3.10 → 3.25 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.734 → 0.771 +0.036identity cos neighbours 0.734 → 0.771d_b rescored +0.674 → +0.688d_a rescored -0.645 → -0.466d_a mined -0.645d_b mined 0.674min_cos_consec (site) 0.8719min_cos_anchor (site) 0.9072dataset emolialang zhspeaker ZH_B00000_S07841total 45.1schain gain +1.1 dBseam step 1.1 dBcrossfades 150/100/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady
(measured, monologue, formal) 第一个事件是内战,他挫折了耕种,打断了商业必然,使谷物价格远远超出收成情况。所会造成的程度,它必然对王国的所有不同市场均或多或少产生了这种影响。尤其是对伦敦附近的市场。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 0.4/6; vocal-burst blend 3.8/10; 17.8s, ZH.
ZH_B00000_S07841_W000021 · in -19.0 dBFS · gain -1.0 dB · emolia-00021
(normal-paced, formal, monologue) 他们要求从最远的地方得到供应。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; very good recording, no background noise; genuineness 0.0/6; vocal-burst blend 3.5/10; 3.0s, ZH.
ZH_B00000_S07841_W000022 · in -18.7 dBFS · gain -1.3 dB · emolia-00021
(measured, monologue, formal) 这两年的高价比两磅十仙令一六三七年以前十六年的平均价格超过三磅五仙令。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 0.6/6; vocal-burst blend 2.0/10; 8.8s, ZH.
ZH_B00000_S07841_W000023 · in -20.3 dBFS · gain +0.3 dB · emolia-00021
(measured, monologue, didactic) 将这个金额分摊在上世纪最后六十四年,中单是这一项,差不多就可以说明为什么在这些年中价格略有上升。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 0.8/6; vocal-burst blend 3.4/10; 10.0s, ZH.
ZH_B00000_S07841_W000024 · in -19.0 dBFS · gain -1.0 dB · emolia-00021
(measured, formal, monologue) 可是,这些虽是最高的价格,却绝不是似乎由内战造成的唯一高价。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 3.0/10; 6.1s, ZH.
ZH_B00000_S07841_W000025 · in -18.1 dBFS · gain -1.9 dB · emolia-00021
Fear ↓  /  Infatuationidentity −0.06 emotion 82 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #5

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation barely there — 0.20, lower than 80 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.65.

At the same time Fear goes the other way, from 0.74 (higher than 74 % of clips in this corpus) to 0.03 (lower than 97 % of clips in this corpus), a change of -0.71. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.25, then +0.15, then +0.01 — a plateau around step 4, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.86 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 49 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.867 before conversion and 0.809 after — it fell by 0.058. Neighbour-to-neighbour the worst pair went 0.768 → 0.845. (The earlier render, with segment 1 left raw, scores 0.585 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.646 in the original and +0.531 after conversion — 82 % of the delta retained, which is most of it. On the other named axis, Fear, -0.710 became -0.394.

Quality. Mean predicted overall quality across the segments went 2.95 → 3.05 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.867 → 0.809 -0.058identity cos neighbours 0.768 → 0.845d_b rescored +0.646 → +0.531d_a rescored -0.710 → -0.394d_a mined -0.710d_b mined 0.646min_cos_consec (site) 0.8608min_cos_anchor (site) 0.9319dataset emolialang enspeaker EN_LgJ4jWtO_oYtotal 47.9schain gain +0.3 dBseam step 2.9 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fairly steady, formal, newsreading) The Maghreb also had far greater known wealth than the rest of Africa, and its location near the entrance to the Mediterranean gave it strategic importance
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.9s, EN.
EN_LgJ4jWtO_oY_W000133 · in -14.8 dBFS · gain -5.2 dB · emolia-01980
(fairly steady, formal, authoritative) France showed a strong interest in Morocco as early as 1830
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 2.0/10; 3.8s, EN.
EN_LgJ4jWtO_oY_W000134 · in -14.0 dBFS · gain -6.0 dB · emolia-01980
(fairly steady, formal, newsreading) The Alaouite dynasty succeeded in maintaining the independence of Morocco in the 18th and 19th centuries, while other states in the region succumbed to Ottoman, French, or British domination
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.4s, EN.
EN_LgJ4jWtO_oY_W000135 · in -14.7 dBFS · gain -5.3 dB · emolia-01980
(steady, newsreading, formal) In the latter part of the 19th century Morocco's instability resulted in European countries intervening to protect investments and to demand economic concessions
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 10.6s, EN.
EN_LgJ4jWtO_oY_W000136 · in -14.3 dBFS · gain -5.7 dB · emolia-01980
(steady, newsreading, formal) The first years of the 20th century saw major diplomatic efforts by European powers, especially France, to further its interests in the region.In the 1890s, the French administration and military in Algiers called for the annexation of the Tuat
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 16.0s, EN.
EN_LgJ4jWtO_oY_W000137 · in -14.7 dBFS · gain -5.3 dB · emolia-01980
Contemplation ↓  /  Painidentity +0.06 emotion REVERSED   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #6

This chain comes from the proxy rule: the same two-sided test as above, but because Pain is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Pain below average — 0.35, lower than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.61.

At the same time Contemplation goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.23 (lower than 77 % of clips in this corpus), a change of -0.72. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.17, then +0.13, then +0.11 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 40 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.759 before conversion and 0.816 after — it rose by 0.057. Neighbour-to-neighbour the worst pair went 0.830 → 0.847. (The earlier render, with segment 1 left raw, scores 0.799 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Pain moved +0.608 in the original and -0.321 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Contemplation, -0.720 became -0.572.

Quality. Mean predicted overall quality across the segments went 3.19 → 3.27 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.759 → 0.816 +0.057identity cos neighbours 0.830 → 0.847d_b rescored +0.608 → -0.321d_a rescored -0.720 → -0.572d_a mined -0.720d_b mined 0.608min_cos_consec (site) 0.8283min_cos_anchor (site) 0.8417dataset emolialang zhspeaker ZH_B00053_S01162total 38.3schain gain +1.4 dBseam step 2.0 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, measured, normally alert, slightly relaxed, fairly steady
(contemplation · monologue, narration) 书虫身为调酒师,这酒量自然是不会差的。几轮下来,李李就醉得不行了,整个人都醉得趴在他身上起不来了。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation; style: monologue, narration; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 3.6/10; 10.3s, ZH.
ZH_B00053_S01162_W000020 · in -17.8 dBFS · gain -2.2 dB · emolia-03804
(monologue, narration) 书虫看着这个迷人的小妞,此刻就躺在自己的身上,那柔软的香醇就在面前,自然是忍不住想要去亲吻他。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 3.8/10; 11.1s, ZH.
ZH_B00053_S01162_W000021 · in -17.4 dBFS · gain -2.6 dB · emolia-03804
(formal, narration) 就在两个人的双唇快碰到一起的时候,书虫突然感觉到腹部一阵的绞痛。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.3/10; 7.4s, ZH.
ZH_B00053_S01162_W000022 · in -17.1 dBFS · gain -3.0 dB · emolia-03804
(narration, authoritative) 连忙推翻压在身上的黎丽,大叫着从沙发上滚落到地上。此刻。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, authoritative; average recording, no background noise; genuineness 1.5/6; vocal-burst blend 2.6/10; 6.9s, ZH.
ZH_B00053_S01162_W000023 · in -18.6 dBFS · gain -1.4 dB · emolia-03804
(pain · formal, monologue) 被他推开的黎丽也迷迷糊糊的醒了过来。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: formal, monologue; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 4.0/10; 3.4s, ZH.
ZH_B00053_S01162_W000024 · in -17.8 dBFS · gain -2.2 dB · emolia-03804
Disgust ↓  /  Concentrationidentity +0.09 emotion 111 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #7

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration below average — 0.29, lower than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.65.

At the same time Disgust goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.33 (lower than 67 % of clips in this corpus), a change of -0.61. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.07, then +0.17, then +0.23 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.71 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 61 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.736 before conversion and 0.831 after — it rose by 0.095. Neighbour-to-neighbour the worst pair went 0.733 → 0.778. (The earlier render, with segment 1 left raw, scores 0.753 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.651 in the original and +0.723 after conversion — 111 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Disgust, -0.605 became -0.120.

Quality. Mean predicted overall quality across the segments went 3.06 → 3.24 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.736 → 0.831 +0.095identity cos neighbours 0.733 → 0.778d_b rescored +0.651 → +0.723d_a rescored -0.605 → -0.120d_a mined -0.605d_b mined 0.651min_cos_consec (site) 0.7069min_cos_anchor (site) 0.7938dataset emolialang zhspeaker ZH_B00036_S07907total 59.7schain gain +3.5 dBseam step 2.0 dBcrossfades 100/100/150/150 ms
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, normally alert, moderate pitch range, light breath
(disgust, intoxication altered states of consciousness · measured, slightly relaxed, fairly steady, conversational) 他他那城市老一般的男朋友is pretty glumsy,什么意思啊?笨拙的on the farm啊,在。 (ahem)
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, intoxication altered states of consciousness; style: conversational, casual; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 1.1/10; 7.2s, ZH.
ZH_B00036_S07907_W000033 · in -21.2 dBFS · gain +1.2 dB · emolia-03639
(confusion, intoxication altered states of consciousness, interest · normal-paced, neutral tension, moderately variable, casual) 好,接下来再看一个look at that full city (ahem) sliker来看那个fool啊,那个傻瓜一样的city sliker怎么着或者好,yes, no idea how to see pegs啊,它是没有。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as confusion, intoxication altered states of consciousness, interest; style: casual, conversational; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 7.3/10; 12.1s, ZH.
ZH_B00036_S07907_W000034 · in -18.4 dBFS · gain -1.6 dB · emolia-03639
(sexual lust, relief, intoxication altered states of consciousness · normal-paced, slightly relaxed, fairly steady, casual) (ahem) 办法去啊这个他也不知道该怎么去饲养这些啊呃猪的啊。好,那么接下来咱们来复习一下今天所讲的内容,总结一下heavy (ahem) (low mumble) (ahem) heater什么意思啊啊,大人物。 (low mumble)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sexual lust, relief, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 5.9/10; 12.8s, ZH.
ZH_B00036_S07907_W000035 · in -21.0 dBFS · gain +1.0 dB · emolia-03639
(measured, slightly relaxed, fairly steady, didactic) Keep a weether right out保持警惕。 City (ahem) (ahem) (ahem) (ahem) sldiger啊表示圆滑事故啊,表示不是特别可靠的这种城市里边的啊这种啊illy sldiger城市网。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 1.9/10; 11.8s, ZH.
ZH_B00036_S07907_W000036 · in -22.7 dBFS · gain +2.7 dB · emolia-03639
(concentration, interest, contemplation · brisk, slightly relaxed, fairly steady, monologue) (surprised gasp) 就是这样。好各位。那么如果大家呢想跟着我们一起练习刚才的重点句的话,那么欢迎大家进入到我们的啊练习社群之中来。那么入群方式呢,就是大家添加屏幕上的客服助角的微信,由他来去引导你入群好。那么今天咱们就先说到这儿,感谢大家,我们下次再见见。
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, interest, contemplation; style: monologue, authoritative; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 6.3/10; 16.6s, ZH.
ZH_B00036_S07907_W000037 · in -22.8 dBFS · gain +2.8 dB · emolia-03639
Fear ↓  /  Thankfulness Gratitudeidentity +0.20 emotion 85 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #8

This chain comes from the proxy rule: the same two-sided test as above, but because Thankfulness Gratitude is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Thankfulness Gratitude below average — 0.33, lower than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.66.

At the same time Fear goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.06 (lower than 94 % of clips in this corpus), a change of -0.67. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.24, then -0.08, then +0.25 — not a clean run: step 3 moves back the other way by 0.08 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.71 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.68 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.71, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 56 s · sv · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.679 before conversion and 0.879 after — it rose by 0.200. Neighbour-to-neighbour the worst pair went 0.750 → 0.879. (The earlier render, with segment 1 left raw, scores 0.734 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.662 in the original and +0.560 after conversion — 85 % of the delta retained, which is most of it. On the other named axis, Fear, -0.675 became -0.435.

Quality. Mean predicted overall quality across the segments went 3.04 → 3.34 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.679 → 0.879 +0.200identity cos neighbours 0.750 → 0.879d_b rescored +0.662 → +0.560d_a rescored -0.675 → -0.435d_a mined -0.666d_b mined 0.662min_cos_consec (site) 0.6791min_cos_anchor (site) 0.7149dataset podcastlang svspeaker 1951total 54.4schain gain +0.4 dBseam step 1.6 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, slightly rough, balanced body, average recording, normally alert, frequent disfluency, somewhat unclear
(measured, slightly relaxed, fairly steady, conversational) hörde jag talas om en också i kommun (low mumble) som ska vara helt jävla fullskät med pengar.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: conversational, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 1.8/10; 6.8s, SV.
1951_00069056 · in -24.7 dBFS · gain +4.7 dB · podcast-01342
(sourness · measured, relaxed, fairly steady, casual) börja jobb igen. Lite grann. (low mumble) Extra då. Och kör lastbil tror jag det var i det här fallet.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as sourness; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 1.6/10; 7.2s, SV.
1951_00071808 · in -21.7 dBFS · gain +1.7 dB · podcast-01329
(teasing, disgust, intoxication altered states of consciousness · normal-paced, neutral tension, moderately variable, casual) man inte kan kosta. Jo, men det är ju, det är ju det. Då pengar är liksom som ett besvär egentligen att du vill inte göra det av med dem.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as teasing, disgust, intoxication altered states of consciousness; style: casual, conversational; average recording, some background noise; genuineness 5.7/6; vocal-burst blend 3.7/10; 13.0s, SV.
1951_00074856 · in -25.3 dBFS · gain +5.3 dB · podcast-01331
(anger, bitterness · measured, neutral tension, fairly steady, casual) Har du kommer till det att du har så mycket pengar så har du ju förmodligen det av den anledningen att du har jobba och sträva. Nu är ju det här helt andra summer, men nu har jag varit lite duktig att spara lite pengar.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, slightly guarded; reads as anger, bitterness; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 5.3/10; 17.6s, SV.
1951_00077992 · in -29.5 dBFS · gain +9.5 dB · podcast-01942
(thankfulness gratitude, affection, infatuation · measured, neutral tension, moderately variable, casual) Jag vill inte ta dem. Nu har jag möjlighet att köpa. Jag har inte köpten planser än till varadsrummet. Nu gjort för länge sen.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as thankfulness gratitude, affection, infatuation; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.8/6; vocal-burst blend 3.3/10; 10.4s, SV.
1951_00080308 · in -25.3 dBFS · gain +5.3 dB · podcast-01940
Astonishment Surprise ↓  /  Reliefidentity −0.06 emotion 92 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #9

This chain comes from the proxy rule: the same two-sided test as above, but because Relief is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Relief below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.63.

At the same time Astonishment Surprise goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.62. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.21, then +0.11, then +0.14 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.54 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.58 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.54, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 36 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.552 before conversion and 0.492 after — it fell by 0.059. Neighbour-to-neighbour the worst pair went 0.552 → 0.541. (The earlier render, with segment 1 left raw, scores 0.509 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.626 in the original and +0.576 after conversion — 92 % of the delta retained, which is essentially all of it. On the other named axis, Astonishment Surprise, -0.620 became -0.419.

Quality. Mean predicted overall quality across the segments went 2.80 → 3.02 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.552 → 0.492 -0.059identity cos neighbours 0.552 → 0.541d_b rescored +0.626 → +0.576d_a rescored -0.620 → -0.419d_a mined -0.620d_b mined 0.626min_cos_consec (site) 0.5832min_cos_anchor (site) 0.5379dataset emolialang zhspeaker ZH_B00066_S03459total 35.2schain gain +3.2 dBseam step 2.9 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a child feminine voice
(astonishment surprise, longing, emotional numbness · measured, energised, slightly relaxed, storytelling) I could see that she was expecting a baby.
full caption & clip details
A child feminine voice; delivery is energised, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as astonishment surprise, longing, emotional numbness; style: storytelling, whispered; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 3.0/10; 3.1s, ZH.
ZH_B00066_S03459_W000493 · in -30.2 dBFS · gain +10.2 dB · emolia-03941
(distress, pain, fear · slow, very low-energy, neutral tension, storytelling) I run all the way here from weaering height, she said, goasping for breath. I couldn't countain many times. I fall en down嗯。
full caption & clip details
A child strongly feminine voice; delivery is very low-energy, slow, neutral tension, variable; timbre is slightly cool, slightly dark, smooth, slightly thin; slurred, frequent disfluency, very wide pitch range, audible breath; affect is negative, submissive, vulnerable; reads as distress, pain, fear; style: storytelling, whispered; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 2.2/10; 11.1s, ZH.
ZH_B00066_S03459_W000494 · in -29.7 dBFS · gain +9.7 dB · emolia-03941
(fear, longing, impatience and irritability · measured, energised, slightly relaxed, storytelling) Please ask me to find some dry clothes for me, and then i'll go onto the village. I'm not staying here.
full caption & clip details
A child strongly feminine voice; delivery is energised, measured, slightly relaxed, moderately variable; timbre is cool, neutral-bright, smooth, thin; clear, almost no disfluency, very wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as fear, longing, impatience and irritability; style: storytelling, whispered; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.2/10; 6.9s, ZH.
ZH_B00066_S03459_W000495 · in -26.6 dBFS · gain +6.6 dB · emolia-03941
(affection, sexual lust, jealousy and envy · measured, normally alert, slightly relaxed, narration) First, my dear young lady, i tota you'll get warm dry, and i'll puts a bandage on that wound. Then we'll have some tea.
full caption & clip details
An elderly feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is cool, slightly dark, slightly rough, thin; clear, almost no disfluency, wide pitch range, audible breath; affect is negative, slightly dominant, slightly guarded; reads as affection, sexual lust, jealousy and envy; style: narration, storytelling; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 0.3/10; 9.7s, ZH.
ZH_B00066_S03459_W000496 · in -29.8 dBFS · gain +9.8 dB · emolia-03941
(relief, affection, thankfulness gratitude · measured, energised, slightly relaxed, narration) She was so exhausted, but she let me help without protesting.
full caption & clip details
An adult feminine voice; delivery is energised, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, almost no disfluency, wide pitch range, light breath; affect is negative, slightly dominant, neutral openness; reads as relief, affection, thankfulness gratitude; style: narration, whispered; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 3.8/10; 5.0s, ZH.
ZH_B00066_S03459_W000497 · in -30.5 dBFS · gain +10.5 dB · emolia-03941
Thankfulness Gratitude ↓  /  Concentrationidentity +0.09 emotion 105 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #10

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.61.

At the same time Thankfulness Gratitude goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.60. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are -0.08, then +0.20, then +0.24, then +0.25 — not a clean run: step 1 moves back the other way by 0.08 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.77 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.71 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.77, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 56 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.561 before conversion and 0.649 after — it rose by 0.088. Neighbour-to-neighbour the worst pair went 0.366 → 0.377. (The earlier render, with segment 1 left raw, scores 0.560 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.609 in the original and +0.637 after conversion — 105 % of the delta retained, which is essentially all of it. On the other named axis, Thankfulness Gratitude, -0.605 became -0.437.

Quality. Mean predicted overall quality across the segments went 2.71 → 2.96 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.561 → 0.649 +0.088identity cos neighbours 0.366 → 0.377d_b rescored +0.609 → +0.637d_a rescored -0.605 → -0.437d_a mined -0.605d_b mined 0.609min_cos_consec (site) 0.7121min_cos_anchor (site) 0.7651dataset emolialang enspeaker EN_WhVmSsF66mItotal 54.5schain gain +4.1 dBseam step 2.6 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · fairly smooth, normal-paced, slightly relaxed, moderate pitch range
(thankfulness gratitude, affection, contentment · normally alert, fairly steady, some disfluency, playful) (ahem) Three documents that Deanna shared with you all today. She gave, (low mumble) uhm, a list of
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, slightly bright, fairly smooth, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as thankfulness gratitude, affection, contentment; style: playful, conversational; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 3.1/10; 5.3s, EN.
EN_WhVmSsF66mI_W000490 · in -17.2 dBFS · gain -2.8 dB · emolia-01783
(hope enthusiasm optimism, contentment, elation · normally alert, fairly steady, some disfluency, casual) You can learn how to caption videos. And of course, the distance learning team is here to help you. And so we're available to all of you so that if you want to caption your own videos, we can help you with that. (ahem) But what I will say is captioning videos is important. If you do take a video from YouTube and it has words and you feel like it's (ahem) correctly captioned,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism, contentment, elation; style: casual, monologue; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 3.7/10; 22.9s, EN.
EN_WhVmSsF66mI_W000492 · in -19.5 dBFS · gain -0.5 dB · emolia-01783
(energised, moderately variable, some disfluency, casual) Please make sure that it is when you're looking for videos because
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: casual, dramatic; below-average recording, quiet background; genuineness 3.3/6; vocal-burst blend 4.3/10; 3.8s, EN.
EN_WhVmSsF66mI_W000493 · in -14.7 dBFS · gain -5.3 dB · emolia-01783
(doubt · normally alert, fairly steady, little disfluency, monologue) Just because YouTube automatically captions it, doesn't mean it's correctly captioned.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as doubt; style: monologue, formal; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.9/10; 4.7s, EN.
EN_WhVmSsF66mI_W000494 · in -18.9 dBFS · gain -1.1 dB · emolia-01783
(concentration · normally alert, fairly steady, some disfluency, casual) (ahem) Uhm, you would know if it's correctly captioned if the video doesn't start automatically. (ahem) Uhm, if the video has periods, commas, has uppercase lettering, uhm, (low mumble) does, does not have words misspelled, that would be correct captioning. And so, (ahem) uhm,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as concentration; style: casual, conversational; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 2.4/10; 18.5s, EN.
EN_WhVmSsF66mI_W000495 · in -18.3 dBFS · gain -1.7 dB · emolia-01783
Concentration ↓  /  Emotional Numbnessidentity +0.27 emotion 88 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #11

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness below average — 0.32, lower than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.62.

At the same time Concentration goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.16 (lower than 84 % of clips in this corpus), a change of -0.77. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.20, then +0.21, then +0.03 — a plateau around step 4, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.52 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.66 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.52, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 36 s · fr · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.391 before conversion and 0.662 after — it rose by 0.270. Neighbour-to-neighbour the worst pair went 0.592 → 0.656. (The earlier render, with segment 1 left raw, scores 0.478 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.616 in the original and +0.541 after conversion — 88 % of the delta retained, which is most of it. On the other named axis, Concentration, -0.773 became -0.702.

Quality. Mean predicted overall quality across the segments went 2.76 → 2.93 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.391 → 0.662 +0.270identity cos neighbours 0.592 → 0.656d_b rescored +0.616 → +0.541d_a rescored -0.773 → -0.702d_a mined -0.773d_b mined 0.616min_cos_consec (site) 0.6572min_cos_anchor (site) 0.5186dataset emolialang frspeaker FR_l_WSrm9kzxutotal 34.6schain gain +1.5 dBseam step 1.6 dBcrossfades 150/150/100/100 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, light breath
(concentration · normal-paced, normally alert, slightly relaxed, monologue) euh, (low mumble) par contre, avec nous, c'est un peu compliqué. Donc, on l'a bloqué, (surprised gasp) elle se retourne beaucoup, donc on la bloque avec ceci, et, (ahem) euh, pour le moment, ça fonctionne très très bien. C'est vrai qu'il n'est pas adapté pour ce lit.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, didactic; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.5/10; 11.3s, FR.
FR_l_WSrm9kzxu_W000009 · in -19.5 dBFS · gain -0.5 dB · emolia-02843
(relief, contemplation, fatigue exhaustion · measured, subdued, slightly relaxed, whispered) Mais on a essayé de trouver des vis un peu plus longues pour, (ahem) pour le fixer et ça a bien fonctionné.
full caption & clip details
A young adult feminine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; reads as relief, contemplation, fatigue exhaustion; style: whispered, ASMR; good recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.0/10; 7.3s, FR.
FR_l_WSrm9kzxu_W000010 · in -17.6 dBFS · gain -2.4 dB · emolia-02843
(measured, normally alert, slightly relaxed, whispered) Voilà sa petite table de nuit. Donc la veilleuse qui était en haut, on l'a déplacée ici en bas. Et voilà.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: whispered, monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.0/10; 6.5s, FR.
FR_l_WSrm9kzxu_W000011 · in -16.7 dBFS · gain -3.3 dB · emolia-02843
(jealousy and envy, emotional numbness · measured, normally alert, slightly relaxed, whispered) Et puis après, (low mumble) euh, voilà, de ce côté-là, il y a son petit bureau et (ahem) son petit coffre de rangement, là où elle peut mettre ses livres.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; reads as jealousy and envy, emotional numbness; style: whispered, ASMR; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.7/10; 7.2s, FR.
FR_l_WSrm9kzxu_W000012 · in -17.7 dBFS · gain -2.3 dB · emolia-02843
(emotional numbness · slow, very low-energy, relaxed, casual) (low mumble) euh, ces, ces outils de peinture, etc.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, thin; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, submissive, neutral openness; reads as emotional numbness; style: casual, ASMR; poor recording, no background noise; genuineness 4.7/6; vocal-burst blend 2.5/10; 3.0s, FR.
FR_l_WSrm9kzxu_W000013 · in -20.3 dBFS · gain +0.3 dB · emolia-02843
Doubt ↓  /  Malevolence Maliceidentity +0.08 emotion 98 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #12

This chain comes from the proxy rule: the same two-sided test as above, but because Malevolence Malice is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Malevolence Malice essentially absent — 0.06, lower than 94 % of clips in this corpus — and ends with it clearly present at 0.68, higher than 68 % of clips in this corpus. That is a total rise of 0.62.

At the same time Doubt goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.37 (lower than 63 % of clips in this corpus), a change of -0.61. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.14, then +0.08, then +0.24, then +0.16 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 34 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.810 before conversion and 0.890 after — it rose by 0.080. Neighbour-to-neighbour the worst pair went 0.810 → 0.862. (The earlier render, with segment 1 left raw, scores 0.772 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Malevolence Malice moved +0.622 in the original and +0.608 after conversion — 98 % of the delta retained, which is essentially all of it. On the other named axis, Doubt, -0.488 became -0.940.

Quality. Mean predicted overall quality across the segments went 2.90 → 3.07 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.810 → 0.890 +0.080identity cos neighbours 0.810 → 0.862d_b rescored +0.622 → +0.608d_a rescored -0.488 → -0.940d_a mined -0.608d_b mined 0.622min_cos_consec (site) 0.8327min_cos_anchor (site) 0.8327dataset emolialang zhspeaker ZH_B00015_S06047total 32.6schain gain +2.3 dBseam step 1.2 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, normally alert, slightly relaxed, some disfluency
(doubt, confusion, intoxication altered states of consciousness · fast, moderately variable, casual, conversational) (chuckle) 这也是喊话嘛,都是让两边的对吧?同期的选手一起上台,我感觉挺有意思的,看看实力是。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, confusion, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 4.2/10; 9.5s, ZH.
ZH_B00015_S06047_W000026 · in -23.4 dBFS · gain +3.4 dB · emolia-03429
(intoxication altered states of consciousness, confusion, triumph · fast, fairly steady, casual, storytelling) 因为上一把队友是完成了一波四抓啊,而且我相信怎么说呢?前面。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as intoxication altered states of consciousness, confusion, triumph; style: casual, storytelling; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 5.2/10; 5.0s, ZH.
ZH_B00015_S06047_W000027 · in -24.1 dBFS · gain +4.0 dB · emolia-03429
(intoxication altered states of consciousness · fast, fairly steady, conversational, casual) 虽然说他的风格怎么说呢,也是比较偏纯正救人味啊,他不玩那些比较激进一点的ob嘛。对。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as intoxication altered states of consciousness; style: conversational, casual; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 6.2/10; 7.6s, ZH.
ZH_B00015_S06047_W000028 · in -24.8 dBFS · gain +4.8 dB · emolia-03429
(intoxication altered states of consciousness · normal-paced, fairly steady, conversational, casual) 这里是杨洋徐俊妍得分六千一百九十二点零七,均前场五十三点八每秒。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as intoxication altered states of consciousness; style: conversational, casual; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 4.7/10; 5.5s, ZH.
ZH_B00015_S06047_W000029 · in -27.6 dBFS · gain +7.6 dB · emolia-03429
(fast, fairly steady, storytelling, casual) 而且上一局打完之后呢,有个什么问题啊,ACD是把先知送上了全局竞选。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: storytelling, casual; average recording, no background noise; genuineness 3.3/6; vocal-burst blend 6.8/10; 5.7s, ZH.
ZH_B00015_S06047_W000030 · in -27.6 dBFS · gain +7.6 dB · emolia-03429
Disgust ↓  /  Contemplationidentity −0.01 emotion 107 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #13

This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contemplation below average — 0.28, lower than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.64.

At the same time Disgust goes the other way, from 0.74 (higher than 74 % of clips in this corpus) to 0.14 (lower than 86 % of clips in this corpus), a change of -0.60. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.23, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 30 s · ko · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.883 before conversion and 0.874 after — it fell by 0.009. Neighbour-to-neighbour the worst pair went 0.891 → 0.859. (The earlier render, with segment 1 left raw, scores 0.770 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.642 in the original and +0.690 after conversion — 107 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Disgust, -0.600 became +0.000.

Quality. Mean predicted overall quality across the segments went 3.00 → 3.21 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.883 → 0.874 -0.009identity cos neighbours 0.891 → 0.859d_b rescored +0.642 → +0.690d_a rescored -0.600 → +0.000d_a mined -0.600d_b mined 0.642min_cos_consec (site) 0.8834min_cos_anchor (site) 0.8834dataset emolialang kospeaker KO_W86C06i3VxYtotal 28.7schain gain +2.0 dBseam step 1.3 dBcrossfades 150/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(brisk, little disfluency, clear, authoritative) 또한, hdmi-264라든지, 그런 비디오 코덱 또한 gpu를 통해서 바로 디코딩이 될 수가 있습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, dramatic; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 2.5/10; 5.1s, KO.
KO_W86C06i3VxY_W000027 · in -20.4 dBFS · gain +0.4 dB · emolia-03120
(normal-paced, some disfluency, average clarity, conversational) 하지만 이쪽까지는 이번 세션의 주요한 목적은 아니고요. 이쪽에 관해서 궁금하신 분들은 앞에서 설명드렸던
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational, casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 3.3/10; 7.7s, KO.
KO_W86C06i3VxY_W000028 · in -20.0 dBFS · gain -0.0 dB · emolia-03120
(brisk, some disfluency, average clarity, authoritative) 그 볼, 그 앞에서 적혀있던 그런 링크들을 통해서 여러 가지 메터리얼이 있으니까 그걸 보시면 될 것 같습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, monologue; good recording, no background noise; genuineness 2.4/6; vocal-burst blend 3.0/10; 5.8s, KO.
KO_W86C06i3VxY_W000029 · in -23.3 dBFS · gain +3.3 dB · emolia-03120
(contemplation · normal-paced, some disfluency, average clarity, conversational) 이렇게 엑셀트 컴포지팅이 다 좋은데 이게 여전히 문제가 있어요. 왜냐면 여전히 느리다는 거죠. 특히 기술이 발전하고 컨텐츠들이 좋아지면 좋아질수록
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contemplation; style: conversational, didactic; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 6.7/10; 10.6s, KO.
KO_W86C06i3VxY_W000030 · in -24.1 dBFS · gain +4.0 dB · emolia-03120
Interest ↓  /  Fatigue Exhaustionidentity −0.03 emotion 90 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #14

This chain comes from the proxy rule: the same two-sided test as above, but because Fatigue Exhaustion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fatigue Exhaustion essentially absent — 0.04, lower than 96 % of clips in this corpus — and ends with it clearly present at 0.73, higher than 73 % of clips in this corpus. That is a total rise of 0.69.

At the same time Interest goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.38 (lower than 62 % of clips in this corpus), a change of -0.62. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are -0.02, then +0.23, then +0.24, then +0.23 — not a clean run: step 1 moves back the other way by 0.02 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.82 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 50 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.611 before conversion and 0.578 after — it fell by 0.033. Neighbour-to-neighbour the worst pair went 0.785 → 0.704. (The earlier render, with segment 1 left raw, scores 0.483 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.695 in the original and +0.628 after conversion — 90 % of the delta retained, which is essentially all of it. On the other named axis, Interest, -0.616 became -0.601.

Quality. Mean predicted overall quality across the segments went 2.93 → 3.05 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.611 → 0.578 -0.033identity cos neighbours 0.785 → 0.704d_b rescored +0.695 → +0.628d_a rescored -0.616 → -0.601d_a mined -0.615d_b mined 0.692min_cos_consec (site) 0.8152min_cos_anchor (site) 0.7907dataset emolialang enspeaker EN_fkddVM6z9qototal 48.5schain gain +1.8 dBseam step 1.7 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, slightly relaxed, fairly steady, some disfluency
(interest, elation, hope enthusiasm optimism · very low-energy, average clarity, casual, conversational) And we think about our, our media relations toolkit or toolbox. (ahem) Um, there could be things in there, including ANRs, (low mumble) uh, op-eds, media kits, satellite media tours, boiler plates, online newsrooms, (ahem) uh, social media, just all of these things, blogs, and just, (ahem) (ahem) uh, lots of different tools.
full caption & clip details
An adult masculine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, elation, hope enthusiasm optimism; style: casual, conversational; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 6.0/10; 17.4s, EN.
EN_fkddVM6z9qo_W000004 · in -18.8 dBFS · gain -1.2 dB · emolia-00330
(concentration, hope enthusiasm optimism, interest · normally alert, somewhat unclear, casual, didactic) that we can utilize and that we can incorporate into (low mumble) our, our media relations toolkit. And you can certainly find a lot of those tools and should find and develop a lot of those tools so that you have them in your tool belt. More tools you have in your tool belt, the better off you will be.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, hope enthusiasm optimism, interest; style: casual, didactic; good recording, quiet background; genuineness 2.6/6; vocal-burst blend 5.0/10; 15.1s, EN.
EN_fkddVM6z9qo_W000005 · in -20.6 dBFS · gain +0.6 dB · emolia-00330
(normally alert, average clarity, casual, monologue) Of course, you need to know how to use them properly, though, and, and when to use them, when to pull out this tool or that tool, and to have the best desired effect.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 2.6/6; vocal-burst blend 3.6/10; 8.6s, EN.
EN_fkddVM6z9qo_W000006 · in -18.6 dBFS · gain -1.4 dB · emolia-00330
(normally alert, average clarity, monologue, formal) I'd like to focus a few minutes then on some of those key considerations rather than the specific tools.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 0.9/10; 4.5s, EN.
EN_fkddVM6z9qo_W000007 · in -16.1 dBFS · gain -3.9 dB · emolia-00330
(normally alert, average clarity, casual, monologue) which you can find in lots, and those are constantly changing anyway, the different specific tools.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.4/6; vocal-burst blend 1.6/10; 3.7s, EN.
EN_fkddVM6z9qo_W000008 · in -17.4 dBFS · gain -2.6 dB · emolia-00330
Disgust ↓  /  Embarrassmentidentity +0.38 emotion 85 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #15

This chain comes from the proxy rule: the same two-sided test as above, but because Embarrassment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Embarrassment below average — 0.30, lower than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.64.

At the same time Disgust goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.33 (lower than 67 % of clips in this corpus), a change of -0.65. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.25, then +0.16, then +0.23 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.17 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.21 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.17, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 23 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.200 before conversion and 0.575 after — it rose by 0.376. Neighbour-to-neighbour the worst pair went 0.205 → 0.590. (The earlier render, with segment 1 left raw, scores 0.516 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.640 in the original and +0.542 after conversion — 85 % of the delta retained, which is most of it. On the other named axis, Disgust, -0.651 became -0.792.

Quality. Mean predicted overall quality across the segments went 2.75 → 2.86 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.200 → 0.575 +0.376identity cos neighbours 0.205 → 0.590d_b rescored +0.640 → +0.542d_a rescored -0.651 → -0.792d_a mined -0.651d_b mined 0.640min_cos_consec (site) 0.2098min_cos_anchor (site) 0.1721dataset emolialang enspeaker EN_vKhSs7ZdFWAtotal 21.8schain gain +1.1 dBseam step 2.9 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · average recording, light breath
(disgust, contempt, impatience and irritability · measured, subdued, slightly relaxed, monologue) I mean, these people present themselves. I mean, officers are killing the line of duty, trying to do exactly what Officer Reed and those officers were doing.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, little disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, contempt, impatience and irritability; style: monologue, formal; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.3/10; 9.0s, EN.
EN_vKhSs7ZdFWA_W000014 · in -20.6 dBFS · gain +0.6 dB · emolia-01995
(doubt · measured, energised, slightly relaxed, authoritative) Are they asking legitimate questions?
full caption & clip details
An adult masculine voice; delivery is energised, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, rough, thin; clear, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as doubt; style: authoritative, formal; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.1/10; 3.3s, EN.
EN_vKhSs7ZdFWA_W000015 · in -17.2 dBFS · gain -2.8 dB · emolia-01995
(confusion, contempt, impatience and irritability · normal-paced, normally alert, slightly relaxed, casual) Oftentimes there are, there are legitimate questions.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as confusion, contempt, impatience and irritability; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 1.6/10; 3.1s, EN.
EN_vKhSs7ZdFWA_W000016 · in -16.3 dBFS · gain -3.7 dB · emolia-01995
(embarrassment, fatigue exhaustion, intoxication altered states of consciousness · normal-paced, normally alert, neutral tension, casual) I'm, (low mumble) uh, okay with questions from officers and my, my auditing style has changed throughout the years.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as embarrassment, fatigue exhaustion, intoxication altered states of consciousness; style: casual, conversational; average recording, some background noise; genuineness 4.0/6; vocal-burst blend 1.3/10; 7.0s, EN.
EN_vKhSs7ZdFWA_W000017 · in -18.9 dBFS · gain -1.1 dB · emolia-01995
Pain ↓  /  Interestidentity +0.33 emotion 112 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #16

This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Interest below average — 0.28, lower than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.69.

At the same time Pain goes the other way, from 0.88 (higher than 88 % of clips in this corpus) to 0.11 (lower than 89 % of clips in this corpus), a change of -0.77. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.25, then +0.18, then +0.06 — a plateau around step 4, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.04 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.08 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.04, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 52 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.048 before conversion and 0.280 after — it rose by 0.328. Neighbour-to-neighbour the worst pair went -0.091 → 0.448. (The earlier render, with segment 1 left raw, scores 0.302 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.691 in the original and +0.772 after conversion — 112 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Pain, -0.773 became -0.727.

Quality. Mean predicted overall quality across the segments went 2.73 → 3.03 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 -0.048 → 0.280 +0.328identity cos neighbours -0.091 → 0.448d_b rescored +0.691 → +0.772d_a rescored -0.773 → -0.727d_a mined -0.773d_b mined 0.691min_cos_consec (site) -0.0789min_cos_anchor (site) -0.0413dataset emolialang enspeaker EN_CCCR7wBD4hItotal 50.4schain gain +2.8 dBseam step 0.9 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · fairly smooth, normally alert, fairly steady, light breath
(normal-paced, slightly relaxed, little disfluency, formal) And although in the plan there was more than 15 confirmed COVID-19 patients.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 1.6/10; 4.5s, EN.
EN_CCCR7wBD4hI_W000372 · in -17.4 dBFS · gain -2.6 dB · emolia-02558
(thankfulness gratitude, confusion · measured, slightly relaxed, little disfluency, monologue) So wearing a mask is a huge difference in this enclosed space. Thank you.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as thankfulness gratitude, confusion; style: monologue, formal; average recording, no background noise; genuineness 1.8/6; vocal-burst blend 1.8/10; 4.7s, EN.
EN_CCCR7wBD4hI_W000373 · in -19.7 dBFS · gain -0.3 dB · emolia-02558
(normal-paced, slightly relaxed, some disfluency) (ahem) Uhm, related to building design, so how do spaces like cubical, (ahem) uh, cubicals compare to new open plan office areas, uh, (ahem) for the resistance to spread of aerosols?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 1.3/10; 12.4s, EN.
EN_CCCR7wBD4hI_W000375 · in -20.5 dBFS · gain +0.5 dB · emolia-02558
(concentration, doubt, contemplation · normal-paced, neutral tension, some disfluency, casual) (ahem) So how we did, we were talking earlier about how we configure buildings. For example, if we have really small cubicles compared to open areas, how would that affect the spread of the virus? (ahem) I think (ahem) Jan, you might be able to take this question as well.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is slightly cool, slightly dark, fairly smooth, slightly thin; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, neutral openness; reads as concentration, doubt, contemplation; style: casual, conversational; below-average recording, some background noise; genuineness 3.5/6; vocal-burst blend 5.0/10; 17.4s, EN.
EN_CCCR7wBD4hI_W000376 · in -19.5 dBFS · gain -0.5 dB · emolia-02558
(interest · brisk, slightly relaxed, some disfluency, casual) (ahem) Depends on what type of ventilation systems you have, right? If you design like modern buildings with the underfloor air distribution, so then (ahem) the cubicle is very good because you supply the air under the floor.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest; style: casual, monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 5.7/10; 12.2s, EN.
EN_CCCR7wBD4hI_W000378 · in -19.2 dBFS · gain -0.8 dB · emolia-02558
Pain ↓  /  Infatuationidentity +0.01 emotion 99 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #17

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation barely there — 0.23, lower than 77 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.62.

At the same time Pain goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.35 (lower than 65 % of clips in this corpus), a change of -0.60. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.25, then +0.20, then -0.05 — not a clean run: step 4 moves back the other way by 0.05 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 34 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.817 before conversion and 0.823 after — it rose by 0.006. Neighbour-to-neighbour the worst pair went 0.853 → 0.868. (The earlier render, with segment 1 left raw, scores 0.769 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.617 in the original and +0.608 after conversion — 99 % of the delta retained, which is essentially all of it. On the other named axis, Pain, -0.602 became -0.708.

Quality. Mean predicted overall quality across the segments went 3.07 → 3.18 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.817 → 0.823 +0.006identity cos neighbours 0.853 → 0.868d_b rescored +0.617 → +0.608d_a rescored -0.602 → -0.708d_a mined -0.602d_b mined 0.617min_cos_consec (site) 0.8813min_cos_anchor (site) 0.9344dataset emolialang zhspeaker ZH_B00064_S03304total 32.5schain gain +0.9 dBseam step 1.5 dBcrossfades 100/100/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, measured, normally alert, slightly relaxed
(pain · monologue, formal) 只关心土豪书做的能量增幅,药水荒歌牌八十二年雪碧能不能用?而此时,科学家地中海告诉白毛书荒歌牌八十二年雪碧有副作用。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: monologue, formal; average recording, no background noise; genuineness 0.2/6; vocal-burst blend 1.8/10; 11.8s, ZH.
ZH_B00064_S03304_W000008 · in -23.3 dBFS · gain +3.3 dB · emolia-03916
(monologue, narration) 会让人变得狂暴。白毛叔一听瞬间怒火攻了心,当即表示,如果不行的话,就不给土豪叔投钱了。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; average recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.9/10; 8.0s, ZH.
ZH_B00064_S03304_W000009 · in -23.1 dBFS · gain +3.0 dB · emolia-03916
(formal, monologue) 为了拿到钱,土豪,叔决定拿自己做实验。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 2.8/10; 3.3s, ZH.
ZH_B00064_S03304_W000010 · in -22.0 dBFS · gain +2.0 dB · emolia-03916
(monologue, narration) 只见土豪叔喝了一瓶荒歌牌,八十二年雪碧,又闻了一阵浓烟之后。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.4/10; 5.7s, ZH.
ZH_B00064_S03304_W000011 · in -22.7 dBFS · gain +2.7 dB · emolia-03916
(monologue, formal) 土豪叔变得十分的狂暴,一耳巴子就把地中海给结了果。
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.5/10; 4.4s, ZH.
ZH_B00064_S03304_W000012 · in -24.2 dBFS · gain +4.2 dB · emolia-03916
Confusion ↓  /  Triumphidentity −0.00 emotion 111 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #18

This chain comes from the proxy rule: the same two-sided test as above, but because Triumph is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Triumph below average — 0.38, lower than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.61.

At the same time Confusion goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.34 (lower than 66 % of clips in this corpus), a change of -0.65. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.20, then +0.12, then +0.07 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.55 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.60 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.55, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 33 s · de · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.636 before conversion and 0.634 after — it fell by 0.002. Neighbour-to-neighbour the worst pair went 0.691 → 0.748. (The earlier render, with segment 1 left raw, scores 0.633 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Triumph moved +0.611 in the original and +0.676 after conversion — 111 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Confusion, -0.650 became -0.347.

Quality. Mean predicted overall quality across the segments went 2.80 → 3.06 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.636 → 0.634 -0.002identity cos neighbours 0.691 → 0.748d_b rescored +0.611 → +0.676d_a rescored -0.650 → -0.347d_a mined -0.650d_b mined 0.611min_cos_consec (site) 0.6026min_cos_anchor (site) 0.5474dataset emolialang despeaker DE_o4H1G-mk3-Utotal 31.9schain gain +2.1 dBseam step 2.4 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-bright, balanced body, quiet background
(confusion, intoxication altered states of consciousness, doubt · slow, normally alert, slightly relaxed, casual) Oh, ja. Ja, was willst denn du? Ah, so.
full caption & clip details
A young adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, frequent disfluency, wide pitch range, audible breath; affect is neutral, neutral stance, neutral openness; reads as confusion, intoxication altered states of consciousness, doubt; style: casual, conversational; below-average recording, quiet background; genuineness 4.5/6; vocal-burst blend 0.0/10; 3.2s, DE.
DE_o4H1G-mk3-U_W000000 · in -21.3 dBFS · gain +1.3 dB · emolia-00141
(longing · normal-paced, normally alert, slightly relaxed, authoritative) Nee, bis jetzt noch nicht, Manu. Außer ein 87er Grießmann Goldkarte. Heute starten wir mal mit der Eins rein, Jungs.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing; style: authoritative, playful; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.0/10; 7.6s, DE.
DE_o4H1G-mk3-U_W000001 · in -19.4 dBFS · gain -0.6 dB · emolia-00141
(relief · normal-paced, normally alert, slightly relaxed, authoritative) er sieht sogar spielbar aus, ne, sag ich dir, wie's ist. Gib dem Shadow drauf und abfahrt.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as relief; style: authoritative, playful; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.2/10; 5.9s, DE.
DE_o4H1G-mk3-U_W000002 · in -17.2 dBFS · gain -2.8 dB · emolia-00141
(malevolence malice, bitterness, jealousy and envy · normal-paced, energised, neutral tension, authoritative) Or, I mean, because of Sentinel. If the pace is enough for you. But you can even play it, ey. Full 3, gg. Do you want Jordi Alba, Prims?
full caption & clip details
A child masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as malevolence malice, bitterness, jealousy and envy; style: authoritative, casual; below-average recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.3/10; 10.7s, DE.
DE_o4H1G-mk3-U_W000003 · in -20.1 dBFS · gain +0.1 dB · emolia-00141
(triumph, disgust, anger · brisk, highly aroused, tense, cartoonish) Konnte das Spiel 10 Tage früher zocken. Konnte Content bringen. Oha! Digga!
full caption & clip details
An adult masculine voice; delivery is highly aroused, brisk, tense, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; very clear, almost no disfluency, very wide pitch range, normal breath; affect is elated, slightly dominant, guarded; reads as triumph, disgust, anger; style: cartoonish, dramatic; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 0.5/10; 5.3s, DE.
DE_o4H1G-mk3-U_W000004 · in -18.4 dBFS · gain -1.6 dB · emolia-00141
Fear ↓  /  Interestidentity −0.00 emotion 69 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #19

This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Interest barely there — 0.24, lower than 76 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.61.

At the same time Fear goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.23 (lower than 77 % of clips in this corpus), a change of -0.70. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.15, then +0.19, then +0.03 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.70 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.75 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.70, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 35 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.665 before conversion and 0.661 after — it fell by 0.004. Neighbour-to-neighbour the worst pair went 0.718 → 0.619. (The earlier render, with segment 1 left raw, scores 0.597 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.609 in the original and +0.418 after conversion — 69 % of the delta retained. On the other named axis, Fear, -0.698 became -0.758.

Quality. Mean predicted overall quality across the segments went 2.40 → 2.83 (+0.43) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.665 → 0.661 -0.004identity cos neighbours 0.718 → 0.619d_b rescored +0.609 → +0.418d_a rescored -0.698 → -0.758d_a mined -0.698d_b mined 0.608min_cos_consec (site) 0.7488min_cos_anchor (site) 0.6970dataset emolialang enspeaker EN_MnH1_C2uQHAtotal 34.0schain gain +3.8 dBseam step 2.9 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-bright, fairly smooth, balanced body, fairly steady, moderate pitch range
(fear · brisk, energised, slightly relaxed, casual) Not necessarily until it's done completely, but before somebody has it in hand.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as fear; style: casual, authoritative; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 1.4/10; 3.6s, EN.
EN_MnH1_C2uQHA_W000024 · in -17.9 dBFS · gain -2.1 dB · emolia-02564
(doubt, astonishment surprise · normal-paced, normally alert, slightly relaxed, authoritative) (low mumble) Uhm, do you, do you count the amount of time it sat in the backlog? (ahem) Inside your task tracking backlog?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, astonishment surprise; style: authoritative, conversational; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 1.0/10; 5.6s, EN.
EN_MnH1_C2uQHA_W000025 · in -22.8 dBFS · gain +2.8 dB · emolia-02564
(impatience and irritability, confusion, doubt · normal-paced, normally alert, slightly relaxed, authoritative) Why would you or why would you not? These are things, these are just things to keep in mind.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, confusion, doubt; style: authoritative, casual; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 0.0/10; 3.4s, EN.
EN_MnH1_C2uQHA_W000026 · in -19.1 dBFS · gain -0.9 dB · emolia-02564
(doubt, concentration · measured, normally alert, neutral tension, monologue) (ahem) Uh, time between deploys can be useful to see how granular improvements to an app are over time. Do we need to make a lot of changes to existing code? (ahem) Uh, just, you know, before we can deploy this new feature, (low mumble) uhm, (low mumble) uh, or, or can, or can we just, can we just add stuff, add new stuff in?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, concentration; style: monologue; average recording, some background noise; genuineness 2.7/6; vocal-burst blend 1.6/10; 16.5s, EN.
EN_MnH1_C2uQHA_W000027 · in -21.6 dBFS · gain +1.6 dB · emolia-02564
(brisk, normally alert, slightly relaxed, authoritative) Time from bug discovery to deployment of the bug fix can be useful in understanding the team's firefighting capabilities.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, dramatic; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 2.0/10; 5.8s, EN.
EN_MnH1_C2uQHA_W000028 · in -20.5 dBFS · gain +0.5 dB · emolia-02564
Astonishment Surprise ↓  /  Disgustidentity −0.07 emotion 69 %   proxy_spearman__PXR__T0.60__C0.25__INTERNAL · #20

This chain comes from the proxy rule: the same two-sided test as above, but because Disgust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Disgust barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it at the very top of the corpus at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.76.

At the same time Astonishment Surprise goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.08 (lower than 92 % of clips in this corpus), a change of -0.64. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.19, then +0.20, then +0.23, then +0.14 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 48 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.780 before conversion and 0.708 after — it fell by 0.072. Neighbour-to-neighbour the worst pair went 0.802 → 0.759. (The earlier render, with segment 1 left raw, scores 0.510 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disgust moved +0.761 in the original and +0.525 after conversion — 69 % of the delta retained. On the other named axis, Astonishment Surprise, -0.640 became -0.576.

Quality. Mean predicted overall quality across the segments went 2.94 → 3.06 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.780 → 0.708 -0.072identity cos neighbours 0.802 → 0.759d_b rescored +0.761 → +0.525d_a rescored -0.640 → -0.576d_a mined -0.640d_b mined 0.761min_cos_consec (site) 0.9610min_cos_anchor (site) 0.9559dataset emolialang enspeaker EN_wby6d6Ua5Lutotal 46.5schain gain +1.9 dBseam step 1.3 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fairly steady, light breath, formal, authoritative) There is evidence that at least two embassies were sent to the Roman Emperor Augustus by Pandya kings
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.6/10; 5.2s, EN.
EN_wby6d6Ua5Lu_W000046 · in -14.3 dBFS · gain -5.7 dB · emolia-01574
(fairly steady, light breath, formal, authoritative) Potsherds with Tamil writing have also been found in excavations on the Red Sea, suggesting the presence of Tamil merchants there
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.1/10; 7.0s, EN.
EN_wby6d6Ua5Lu_W000047 · in -15.5 dBFS · gain -4.5 dB · emolia-01574
(fairly steady, light breath, formal, newsreading) An anonymous 1st century traveller's account written in Greek, Periplus Maris Arithrae, describes the ports of the Pandya and Shara kingdoms in Damarica and their commercial activity in great detail
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; mildly explicit content; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.7s, EN.
EN_wby6d6Ua5Lu_W000048 · in -15.5 dBFS · gain -4.5 dB · emolia-01574
(emotional numbness · steady, minimal breath, newsreading, formal) Peri Plus also indicates that the chief exports of the ancient Tamils were pepper, malabathrum, pearls, ivory, silk, spikenard, diamonds, sapphires, and tortoise shell.The Classical period ended around the 4th century CE with invasions by the Calabra, referred to as the Calapyrur in Tamil literature and inscriptions.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 18.7s, EN.
EN_wby6d6Ua5Lu_W000049 · in -15.5 dBFS · gain -4.5 dB · emolia-01574
(disgust · fairly steady, light breath, formal, newsreading) These invaders are described as evil kings and barbarians coming from lands to the north of the Tamil country
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.2/10; 5.7s, EN.
EN_wby6d6Ua5Lu_W000050 · in -13.8 dBFS · gain -6.2 dB · emolia-01574