k-AB2-k4 — voice-corrected

AB2 at chain length k=4, all corpora, at the mining floor.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_k-AB2-k4.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
60segments re-voiced
0.755 → 0.782median worst-to-anchor identity cosine
87 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Sexual Lust ↓  /  Fatigue Exhaustionidentity −0.08 emotion REVERSED   k-AB2-k4 · #1

This chain comes from the two-sided rule: it only counts if both emotions move — Sexual Lust down and Fatigue Exhaustion up — by at least 0.25 each.

The chain starts with Fatigue Exhaustion around average — 0.52, higher than 52 % of clips in this corpus — and ends with it strongly present at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.28.

At the same time Sexual Lust goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.63 (higher than 63 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.19, then -0.07, then +0.15 — not a clean run: step 2 moves back the other way by 0.07 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.60 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.56 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.60, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 35 s · ja · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.536 before conversion and 0.455 after — it fell by 0.081. Neighbour-to-neighbour the worst pair went 0.526 → 0.689. (The earlier render, with segment 1 left raw, scores 0.445 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Fatigue Exhaustion moved +0.280 in the original and -0.108 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Sexual Lust, -0.338 became -0.378.

Quality. Mean predicted overall quality across the segments went 2.61 → 3.11 (+0.49) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.536 → 0.455 -0.081identity cos neighbours 0.526 → 0.689d_b rescored +0.280 → -0.108d_a rescored -0.338 → -0.378d_a mined -0.338d_b mined 0.280min_cos_consec (site) 0.5571min_cos_anchor (site) 0.6037dataset emolialang jaspeaker JA_B00004_S09207total 33.6schain gain +0.4 dBseam step 1.8 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, average recording, quiet background, some disfluency, somewhat unclear
(sexual lust, pride, disgust · fast, normally alert, neutral tension, casual) こんな感じのホテルになります。まずね、チェックインをして、ご飯食べてないので、ご飯食べに行こうかなと思ってるんですけども、まあ今回はね、ホテルの紹介なので、ホテルの紹介だけをね、やっていこうと思います。それでは、チェックインしてきまーす。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sexual lust, pride, disgust; style: casual, monologue; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 10.0/10; 13.9s, JA.
JA_B00004_S09207_W000004 · in -15.6 dBFS · gain -4.4 dB · emolia-03000
(disgust, embarrassment · fast, normally alert, neutral tension, casual) (low mumble) ちょっと見えないね。ベーシックラインとね、書いてあるんですけども、ホテルの外観はこんな感じで、コインランドリーがあって、手前が、バーがあって、レンタルバイクもここで借りれると。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, slightly dark, fairly smooth, thin; somewhat unclear, some disfluency, moderate pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as disgust, embarrassment; style: casual, storytelling; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 9.5/10; 12.0s, JA.
JA_B00004_S09207_W000005 · in -16.9 dBFS · gain -3.1 dB · emolia-03000
(measured, very low-energy, slightly relaxed, whispered) (ahem) はい。あと、スムージーですね。こんな感じのメニューになってます。
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, neutral openness; no dominant emotion; style: whispered, monologue; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 5.5/10; 5.3s, JA.
JA_B00004_S09207_W000006 · in -14.6 dBFS · gain -5.5 dB · emolia-03000
(fast, normally alert, slightly relaxed, casual) こんな感じになります。じゃあ、ちょっと開けてみましょう。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 5.7/6; vocal-burst blend 3.5/10; 3.1s, JA.
JA_B00004_S09207_W000007 · in -14.4 dBFS · gain -5.6 dB · emolia-03000
Doubt ↓  /  Concentrationidentity −0.02 emotion 100 %   k-AB2-k4 · #2

This chain comes from the two-sided rule: it only counts if both emotions move — Doubt down and Concentration up — by at least 0.25 each.

The chain starts with Concentration clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.38.

At the same time Doubt goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.25, then +0.14, then -0.01 — not a clean run: step 3 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.97 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 47 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.690 before conversion and 0.671 after — it fell by 0.019. Neighbour-to-neighbour the worst pair went 0.873 → 0.836. (The earlier render, with segment 1 left raw, scores 0.595 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.377 in the original and +0.377 after conversion — 100 % of the delta retained, which is essentially all of it. On the other named axis, Doubt, -0.269 became -0.146.

Quality. Mean predicted overall quality across the segments went 2.97 → 3.15 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.690 → 0.671 -0.019identity cos neighbours 0.873 → 0.836d_b rescored +0.377 → +0.377d_a rescored -0.269 → -0.146d_a mined -0.269d_b mined 0.377min_cos_consec (site) 0.9656min_cos_anchor (site) 0.9589dataset emolialang enspeaker EN_kOuRrntlaZototal 46.4schain gain +1.7 dBseam step 0.9 dBcrossfades 150/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(doubt, malevolence malice · formal, authoritative) If the judge himself wishes to bring an accusation, the superior appoints the judge who is to hear it
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, malevolence malice; style: formal, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.0/10; 5.2s, EN.
EN_kOuRrntlaZo_W000028 · in -14.4 dBFS · gain -5.6 dB · emolia-02339
(disappointment, pain, bitterness · newsreading, formal) The decision of an incompetent judge is valid if by common error – error communis – he is held to be competent in civil disputes the parties can entrust the decision to any desired arbiter
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, pain, bitterness; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 9.9s, EN.
EN_kOuRrntlaZo_W000029 · in -15.5 dBFS · gain -4.5 dB · emolia-02339
(concentration, fear · newsreading, formal) If the judge render a defective decision, appeal can be taken to the next higher judge.This relation of the courts to one another and the successive course of appeals
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, fear; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 12.8s, EN.
EN_kOuRrntlaZo_W000030 · in -14.3 dBFS · gain -5.7 dB · emolia-02339
(concentration · newsreading, formal) From the beginning the bishop, or his representative, the archdeacon, or the official officialis, or the vicar general, was the judge in first instance for all suits, contentious or criminal, which arose in the diocese or in the corresponding administrative district, so far as such suits were not withdrawn from his jurisdiction by the common law.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 19.1s, EN.
EN_kOuRrntlaZo_W000031 · in -15.6 dBFS · gain -4.5 dB · emolia-02339
Impatience and Irritability ↓  /  Prideidentity +0.43 emotion REVERSED   k-AB2-k4 · #3

This chain comes from the two-sided rule: it only counts if both emotions move — Impatience and Irritability down and Pride up — by at least 0.25 each.

The chain starts with Pride clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.26.

At the same time Impatience and Irritability goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.59 (higher than 59 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.12, then +0.06, then +0.08 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.42 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.40 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.42, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 57 s · bg · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.356 before conversion and 0.782 after — it rose by 0.426. Neighbour-to-neighbour the worst pair went 0.357 → 0.844. (The earlier render, with segment 1 left raw, scores 0.542 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

The emotional move did not survive. Re-scored end to end, Pride moved +0.256 in the original and -0.089 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Impatience and Irritability, -0.406 became -0.189.

Quality. Mean predicted overall quality across the segments went 3.03 → 3.23 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.356 → 0.782 +0.426identity cos neighbours 0.357 → 0.844d_b rescored +0.256 → -0.089d_a rescored -0.406 → -0.189d_a mined -0.406d_b mined 0.257min_cos_consec (site) 0.3993min_cos_anchor (site) 0.4208dataset eurospeechlang bgspeaker bulgaria_bulgaria_0_110120total 56.4schain gain +0.9 dBseam step 2.0 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · slightly cool, average recording, quiet background, moderately variable
(impatience and irritability, bitterness, disgust · brisk, energised, neutral tension, dramatic) седмици? Раздадени са основно на общини, които се управляват от кметове от ГЕРБ, и ние разполагаме с конкретни факти. Какво става с обществените поръчки? Спрени ли са или не са спрени?
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, almost no disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as impatience and irritability, bitterness, disgust; style: dramatic, authoritative; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 3.7/10; 12.1s, BG.
bulgaria_bulgaria_0_11012017_1938960_1951088 · in -19.8 dBFS · gain -0.2 dB · eurospeech-00032
(anger, contempt, disgust · brisk, energised, neutral tension, dramatic) Какво става с блокираната пътна мрежа на страната? Госпожо Председател на Народното събрание, свържете се с Премиера, припомнете му, че България е парламентарна република и да не влизаме в прецедент, в който Вие криете правителството от блицконтрола!
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as anger, contempt, disgust; style: dramatic, authoritative; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 6.1/10; 18.7s, BG.
bulgaria_bulgaria_0_11012017_1951088_1969776 · in -22.0 dBFS · gain +2.0 dB · eurospeech-00032
(disgust, malevolence malice, contempt · fast, energised, neutral tension, dramatic) Единственото извинение би било, ако премиерът и неговите заместници в момента ринат финала на магистрала „Хемус“ и търсят в снега лентичката от нейното откриване. Благодаря Ви.
full caption & clip details
A young adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, almost no disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, fairly guarded; reads as disgust, malevolence malice, contempt; style: dramatic, ranting; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 4.9/10; 12.3s, BG.
bulgaria_bulgaria_0_11012017_1969776_1982032 · in -21.2 dBFS · gain +1.2 dB · eurospeech-00032
(pride, doubt, helplessness · measured, very low-energy, relaxed, casual) Уважаеми господин Гечев, видно от проекта за седмичната програма, който съм обявила още във вчерашния ден по обяд,
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is slightly cool, slightly dark, slightly rough, thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as pride, doubt, helplessness; style: casual, monologue; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 1.7/10; 14.0s, BG.
bulgaria_bulgaria_0_11012017_1982032_1995984 · in -26.0 dBFS · gain +6.0 dB · eurospeech-00032
Doubt ↓  /  Contemplationidentity −0.07 emotion 37 %   k-AB2-k4 · #4

This chain comes from the two-sided rule: it only counts if both emotions move — Doubt down and Contemplation up — by at least 0.25 each.

The chain starts with Contemplation around average — 0.51, higher than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.48.

At the same time Doubt goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.35. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.17, then +0.21, then +0.10 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 32 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.829 before conversion and 0.757 after — it fell by 0.071. Neighbour-to-neighbour the worst pair went 0.826 → 0.637. (The earlier render, with segment 1 left raw, scores 0.715 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.480 in the original and +0.178 after conversion — 37 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Doubt, -0.319 became -0.296.

Quality. Mean predicted overall quality across the segments went 2.36 → 2.88 (+0.52) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.829 → 0.757 -0.071identity cos neighbours 0.826 → 0.637d_b rescored +0.480 → +0.178d_a rescored -0.319 → -0.296d_a mined -0.346d_b mined 0.482min_cos_consec (site) 0.8270min_cos_anchor (site) 0.8270dataset emolialang enspeaker EN_smVuabjN6E8total 31.3schain gain +0.7 dBseam step 1.9 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · slightly thin, very low-energy, relaxed, audible breath
(doubt, embarrassment, hope enthusiasm optimism · measured, fairly steady, frequent disfluency, ASMR) better in the future, I might redo some of them, or I might write new quotes in the future, (low mumble) uhm, in which case I might do an update to this video, I don't know.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is slightly cool, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, submissive, slightly vulnerable; reads as doubt, embarrassment, hope enthusiasm optimism; style: ASMR, whispered; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 3.5/10; 8.8s, EN.
EN_smVuabjN6E8_W000083 · in -16.4 dBFS · gain -3.6 dB · emolia-01492
(affection, embarrassment, sexual lust · slow, steady, frequent disfluency, ASMR) Do it. So, you know, don't beat yourself up if you don't immediately come up with something that you're super proud of.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, submissive, slightly vulnerable; reads as affection, embarrassment, sexual lust; style: ASMR, whispered; below-average recording, quiet background; genuineness 3.7/6; vocal-burst blend 2.7/10; 7.0s, EN.
EN_smVuabjN6E8_W000085 · in -16.0 dBFS · gain -4.0 dB · emolia-01492
(sexual lust, fear · slow, steady, frequent disfluency, ASMR) tools to shift your mindset whenever you're feeling anxious or
full caption & clip details
An elderly feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, smooth, slightly thin; slurred, frequent disfluency, narrow pitch range, audible breath; affect is mildly positive, submissive, slightly vulnerable; reads as sexual lust, fear; style: ASMR, whispered; average recording, no background noise; genuineness 2.4/6; vocal-burst blend 2.3/10; 5.9s, EN.
EN_smVuabjN6E8_W000087 · in -16.1 dBFS · gain -3.9 dB · emolia-01492
(contemplation, affection, contentment · slow, moderately variable, some disfluency, ASMR) Anything else, any perception that you want to shift in yourself, if you want to tell yourself, if you want to teach yourself to think of yourself as more beautiful, more loving,
full caption & clip details
An elderly feminine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is slightly warm, dark, slightly rough, slightly thin; somewhat unclear, some disfluency, narrow pitch range, audible breath; affect is mildly positive, submissive, slightly vulnerable; reads as contemplation, affection, contentment; style: ASMR, whispered; below-average recording, quiet background; genuineness 3.6/6; vocal-burst blend 4.3/10; 10.2s, EN.
EN_smVuabjN6E8_W000088 · in -15.3 dBFS · gain -4.7 dB · emolia-01492
Relief ↓  /  Interestidentity −0.04 emotion 148 %   k-AB2-k4 · #5

This chain comes from the two-sided rule: it only counts if both emotions move — Relief down and Interest up — by at least 0.25 each.

The chain starts with Interest clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.25.

At the same time Relief goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.74 (higher than 74 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.11, then +0.14, then -0.00 — not a clean run: step 3 moves back the other way by 0.00 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 56 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.856 before conversion and 0.819 after — it fell by 0.038. Neighbour-to-neighbour the worst pair went 0.879 → 0.872. (The earlier render, with segment 1 left raw, scores 0.760 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.251 in the original and +0.372 after conversion — 148 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Relief, -0.257 became -0.126.

Quality. Mean predicted overall quality across the segments went 3.00 → 3.31 (+0.32) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.856 → 0.819 -0.038identity cos neighbours 0.879 → 0.872d_b rescored +0.251 → +0.372d_a rescored -0.257 → -0.126d_a mined -0.257d_b mined 0.251min_cos_consec (site) 0.8943min_cos_anchor (site) 0.8923dataset emolialang zhspeaker ZH_B00079_S01625total 55.1schain gain +0.5 dBseam step 2.3 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, quiet background, normally alert, some disfluency, average clarity, light breath
(relief, sexual lust, pleasure ecstasy · fast, neutral tension, moderately variable, casual) (ahem) 十年前我就知道这个片子,当时那时候豆瓣儿评分就特别高,九点几分。但是每次我一想说,我要看这个片子的时候,就觉得哎呀,太惨了,真看不下去。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as relief, sexual lust, pleasure ecstasy; style: casual, storytelling; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 8.0/10; 8.7s, ZH.
ZH_B00079_S01625_W000022 · in -18.1 dBFS · gain -1.9 dB · emolia-04067
(amusement, hope enthusiasm optimism, sexual lust · brisk, slightly relaxed, moderately variable, casual) 所以也是借这个机会,逼着自己把这个片子看完了啊。第二天大家一到组委会集合的时候,当时不是每个人说一说自己的看法吗?我当时一愣啊,还得说自己的看法呢。
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as amusement, hope enthusiasm optimism, sexual lust; style: casual, dramatic; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 6.7/10; 10.3s, ZH.
ZH_B00079_S01625_W000023 · in -19.0 dBFS · gain -1.0 dB · emolia-04067
(interest, astonishment surprise, relief · brisk, slightly relaxed, fairly steady, casual) 萧瑟的景观啊,你得选暖色调的,得让大家对你这个项目的第一印象是什么?是什么是什么?我的天,这真是太细了哦。第一次做电影策划,感受到这个流程真的是还是有很多门道的。不只是说你有好故事,就能把你自己的产品卖出去,这距离还是很远的。还有很多细节的操作需要去完整的。
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as interest, astonishment surprise, relief; style: casual, conversational; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 10.0/10; 18.6s, ZH.
ZH_B00079_S01625_W000024 · in -17.8 dBFS · gain -2.2 dB · emolia-04067
(interest, thankfulness gratitude, elation · normal-paced, slightly relaxed, fairly steady, casual) 两个小时的杨超导演的讲座,对吧?杨超导演那个讲座,我自己觉得是有一些收获吧。但是我觉得他回避了一些比较关键的问题啊,因为官方给这个活动的定名叫策略。规划局其实是想说给大家一个坚定信念的过程,就是让大家去。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as interest, thankfulness gratitude, elation; style: casual, conversational; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 10.0/10; 18.1s, ZH.
ZH_B00079_S01625_W000025 · in -17.5 dBFS · gain -2.5 dB · emolia-04067
Hope Enthusiasm Optimism ↓  /  Confusionidentity −0.04 emotion 114 %   k-AB2-k4 · #6

This chain comes from the two-sided rule: it only counts if both emotions move — Hope Enthusiasm Optimism down and Confusion up — by at least 0.25 each.

The chain starts with Confusion clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.26.

At the same time Hope Enthusiasm Optimism goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.62 (higher than 62 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.23, then -0.07, then +0.11 — not a clean run: step 2 moves back the other way by 0.07 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 77 s · en · evasnippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.845 before conversion and 0.809 after — it fell by 0.036. Neighbour-to-neighbour the worst pair went 0.848 → 0.764. (The earlier render, with segment 1 left raw, scores 0.749 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.260 in the original and +0.297 after conversion — 114 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Hope Enthusiasm Optimism, -0.377 became -0.357.

Quality. Mean predicted overall quality across the segments went 3.09 → 3.24 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.845 → 0.809 -0.036identity cos neighbours 0.848 → 0.764d_b rescored +0.260 → +0.297d_a rescored -0.377 → -0.357d_a mined -0.377d_b mined 0.259min_cos_consec (site) 0.8730min_cos_anchor (site) 0.8670dataset evasnippetslang enspeaker cond_podcastt_03_02_0003_4total 76.0schain gain +2.4 dBseam step 1.4 dBcrossfades 100/100/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, balanced body, average recording, quiet background, moderately variable, light breath
(hope enthusiasm optimism, elation, interest · measured, very low-energy, relaxed, casual) Everyone, (low mumble) um we will we'll get into more about who Juliana is, of course, and then we'll jump into topics. But (low mumble) uh welcome, welcome everyone. Um (low mumble) welcome Juliana and super excited to have this episode going today. I'm gonna keep it straightforward and simple because I'm excited to talk about whatever we're gonna talk about because (breathy giggle) we did not really we I actually this is I think the one that I plan the least for.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as hope enthusiasm optimism, elation, interest; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 7.7/10; 26.1s.
cond_podcastt_03_02_0003_47667_00004456 · in -23.6 dBFS · gain +3.6 dB · evasnippets-00162
(embarrassment, intoxication altered states of consciousness, helplessness · normal-paced, normally alert, relaxed, casual) I'm really excited to see what the heck's gonna happen. But anyways, (low mumble) um this is it. I don't have anything housekeeping other than (low mumble) uh to throw it out there for those of you who are following along. This is episode seventeen. We're gonna make it to episode twenty and it's gonna go crazy, but (low mumble) uh we're making it and I'm I'm like getting hyped like for the end of this season. I don't know. Hopefully it's as exciting as I'm making it out to be, but
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as embarrassment, intoxication altered states of consciousness, helplessness; style: casual, conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 8.5/10; 24.2s.
cond_podcastt_03_02_0003_47667_00008352 · in -23.2 dBFS · gain +3.2 dB · evasnippets-00008
(sexual lust, astonishment surprise, interest · normal-paced, normally alert, relaxed, casual) Yeah, have you seen that thing where it's like it shows an animal skeleton and be like what scientists think it looks like and then like how it actually looked? I haven't seen that. Okay, wait, I'm gonna pull it up'cause I think it's a rabbit.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as sexual lust, astonishment surprise, interest; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 10.0/10; 11.4s.
cond_podcastt_03_02_0003_47667_00165056 · in -24.5 dBFS · gain +4.5 dB · evasnippets-00325
(confusion, intoxication altered states of consciousness, interest · normal-paced, normally alert, neutral tension, casual) Which is nuts. Well,'cause there's the one I I don't remember what part of the world it is, but I'm just gonna like guess it was the Amazon. Just'cause it's so vast. But like where they would literally like you've seen the birds, like they dance, like they like put their feathers up on their head and like they put'em behind their back and
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as confusion, intoxication altered states of consciousness, interest; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.3/6; vocal-burst blend 10.0/10; 14.8s.
cond_podcastt_03_02_0003_47667_00214519 · in -23.2 dBFS · gain +3.2 dB · evasnippets-00325
Fear ↓  /  Contemplationidentity −0.04 emotion 47 %   k-AB2-k4 · #7

This chain comes from the two-sided rule: it only counts if both emotions move — Fear down and Contemplation up — by at least 0.25 each.

The chain starts with Contemplation around average — 0.54, higher than 54 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.31.

At the same time Fear goes the other way, from 0.87 (higher than 87 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.35. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.22, then +0.04, then +0.05 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 42 s · de · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.755 before conversion and 0.718 after — it fell by 0.036. Neighbour-to-neighbour the worst pair went 0.680 → 0.607. (The earlier render, with segment 1 left raw, scores 0.539 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.308 in the original and +0.146 after conversion — 47 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Fear, -0.353 became -0.268.

Quality. Mean predicted overall quality across the segments went 2.86 → 3.13 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.755 → 0.718 -0.036identity cos neighbours 0.680 → 0.607d_b rescored +0.308 → +0.146d_a rescored -0.353 → -0.268d_a mined -0.353d_b mined 0.309min_cos_consec (site) 0.8270min_cos_anchor (site) 0.8102dataset emolialang despeaker DE_QmUTeb4z0egtotal 41.2schain gain +3.5 dBseam step 0.6 dBcrossfades 150/150/100 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: an adult feminine voice · neutral-toned, fairly smooth, slightly relaxed, fairly steady, moderate pitch range
(measured, normally alert, some disfluency, didactic) Wenn man mal da die Idee hat, was man hier fassen will, kommt es zur Ideation Phase. Da werden
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: didactic, formal; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.3/10; 7.7s, DE.
DE_QmUTeb4z0eg_W000008 · in -13.8 dBFS · gain -6.2 dB · emolia-00252
(concentration · slow, very low-energy, frequent disfluency, monologue) Da wird in interdisziplinären Teams gearbeitet, werden Ideen erfasst, werden oft Kärtchen geschrieben, verschoben und so die Idee (low mumble) geprägt in die verschiedene Wissensgebiete.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, neutral openness; reads as concentration; style: monologue, didactic; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.0/10; 16.4s, DE.
DE_QmUTeb4z0eg_W000009 · in -14.0 dBFS · gain -6.0 dB · emolia-00252
(interest, concentration, bitterness · normal-paced, normally alert, some disfluency, monologue) (low mumble) Design-Elemente, oft sind da Künstler dabei, auf jeden Fall sind verschiedene Anwender dabei, Informatiker, Designer, Ergonomen und so weiter. Da werden also die Ideen kreiert und festgehalten.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as interest, concentration, bitterness; style: monologue, narration; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.0/10; 13.8s, DE.
DE_QmUTeb4z0eg_W000010 · in -13.5 dBFS · gain -6.5 dB · emolia-00252
(slow, subdued, frequent disfluency, casual) generierung kommt, also, es müssen verschiedene,
full caption & clip details
A young adult feminine voice; delivery is subdued, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly submissive, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 0.1/10; 3.8s, DE.
DE_QmUTeb4z0eg_W000011 · in -12.4 dBFS · gain -7.6 dB · emolia-00252
Interest ↓  /  Pleasure Ecstasyidentity +0.03 emotion 134 %   k-AB2-k4 · #8

This chain comes from the two-sided rule: it only counts if both emotions move — Interest down and Pleasure Ecstasy up — by at least 0.25 each.

The chain starts with Pleasure Ecstasy clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.28.

At the same time Interest goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.63 (higher than 63 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.07, then +0.19, then +0.03 — a plateau around step 3, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 43 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.624 before conversion and 0.653 after — it rose by 0.029. Neighbour-to-neighbour the worst pair went 0.662 → 0.640. (The earlier render, with segment 1 left raw, scores 0.397 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pleasure Ecstasy moved +0.284 in the original and +0.381 after conversion — 134 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Interest, -0.319 became -0.121.

Quality. Mean predicted overall quality across the segments went 2.64 → 2.88 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.624 → 0.653 +0.029identity cos neighbours 0.662 → 0.640d_b rescored +0.284 → +0.381d_a rescored -0.319 → -0.121d_a mined -0.319d_b mined 0.284min_cos_consec (site) 0.8470min_cos_anchor (site) 0.8092dataset podcastlang enspeaker 599178total 42.1schain gain +4.6 dBseam step 0.8 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, normally alert, some disfluency, average clarity, light breath
(interest · normal-paced, slightly relaxed, fairly steady, casual) getting the information about understanding what investing management is and what their options are.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as interest; style: casual, conversational; average recording, no background noise; genuineness 4.1/6; vocal-burst blend 4.6/10; 5.4s, EN.
599178_00183880 · in -24.3 dBFS · gain +4.3 dB · podcast-04609
(impatience and irritability · normal-paced, neutral tension, moderately variable, casual) when you start working with a financial advisor and they ask you to trans, they tell you, here's how you transfer over all your accounts. That is not how it like you don't have to do that. That's something they would love
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as impatience and irritability; style: casual, conversational; good recording, quiet background; genuineness 3.4/6; vocal-burst blend 5.5/10; 11.4s, EN.
599178_00185016 · in -24.4 dBFS · gain +4.4 dB · podcast-04620
(affection, jealousy and envy, thankfulness gratitude · brisk, neutral tension, moderately variable, casual) to manage all of your money, but you get to pick. There's no pushback because I know, and I'm guessing that almost never happens that someone
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as affection, jealousy and envy, thankfulness gratitude; style: casual, conversational; average recording, quiet background; genuineness 5.5/6; vocal-burst blend 8.0/10; 21.1s, EN.
599178_00186152 · in -25.0 dBFS · gain +5.0 dB · podcast-04618
(pleasure ecstasy · normal-paced, slightly relaxed, fairly steady, conversational) that. And I've had so many experiences of being on calls with financial professionals,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as pleasure ecstasy; style: conversational, casual; good recording, quiet background; genuineness 4.4/6; vocal-burst blend 2.4/10; 4.7s, EN.
599178_00188440 · in -25.1 dBFS · gain +5.0 dB · podcast-04614
Embarrassment ↓  /  Impatience and Irritabilityidentity +0.04 emotion REVERSED   k-AB2-k4 · #9

This chain comes from the two-sided rule: it only counts if both emotions move — Embarrassment down and Impatience and Irritability up — by at least 0.25 each.

The chain starts with Impatience and Irritability clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.30.

At the same time Embarrassment goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.19, then +0.09, then +0.01 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.76 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.79 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.76, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 36 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.700 before conversion and 0.741 after — it rose by 0.041. Neighbour-to-neighbour the worst pair went 0.813 → 0.747. (The earlier render, with segment 1 left raw, scores 0.604 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Impatience and Irritability moved +0.298 in the original and -0.040 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Embarrassment, -0.268 became -0.567.

Quality. Mean predicted overall quality across the segments went 2.66 → 3.01 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.700 → 0.741 +0.041identity cos neighbours 0.813 → 0.747d_b rescored +0.298 → -0.040d_a rescored -0.268 → -0.567d_a mined -0.270d_b mined 0.297min_cos_consec (site) 0.7914min_cos_anchor (site) 0.7614dataset emolialang enspeaker EN_B00045_S06886total 34.7schain gain +5.3 dBseam step 1.4 dBcrossfades 100/150/100 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice
(embarrassment, amusement, contemplation · measured, normally alert, neutral tension, storytelling) (ahem) You're right. No, Van, I, well, I was as good with him as I knew how. I explained to him that sometimes when people get married, they have bad luck. And their personalities just don't fit in all, and they're terribly unhappy, and they make everyone around them unhappy.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is warm, dark, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as embarrassment, amusement, contemplation; style: storytelling, conversational; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 5.1/10; 17.6s, EN.
EN_B00045_S06886_W000004 · in -22.0 dBFS · gain +2.0 dB · emolia-01120
(infatuation, confusion, fear · slow, very low-energy, fully relaxed, monologue) Nothing at first, just stared at me. Those little angry eyes just glared at me.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, fully relaxed, steady; timbre is warm, slightly dark, rough, very full; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is negative, slightly dominant, fairly guarded; reads as infatuation, confusion, fear; style: monologue, whispered; below-average recording, quiet background; genuineness 1.3/6; vocal-burst blend 0.0/10; 6.7s, EN.
EN_B00045_S06886_W000005 · in -24.6 dBFS · gain +4.6 dB · emolia-01120
(distress, disappointment, pain · normal-paced, normally alert, neutral tension, ranting) I told him that a divorce was to try to give everybody concerned another chance. What did I get for my pain?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as distress, disappointment, pain; style: ranting, storytelling; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 1.4/10; 6.0s, EN.
EN_B00045_S06886_W000006 · in -21.6 dBFS · gain +1.6 dB · emolia-01120
(impatience and irritability, amusement, teasing · brisk, highly aroused, slightly tense, ranting) You said it was a quitter, okay. You said it was my fault, okay again.
full caption & clip details
An adult masculine voice; delivery is highly aroused, brisk, slightly tense, moderately variable; timbre is slightly cool, neutral-bright, very rough, thin; very clear, almost no disfluency, wide pitch range, normal breath; affect is negative, dominant, guarded; reads as impatience and irritability, amusement, teasing; style: ranting, cartoonish; poor recording, some background noise; mildly explicit content; genuineness 2.2/6; vocal-burst blend 2.3/10; 4.7s, EN.
EN_B00045_S06886_W000007 · in -19.3 dBFS · gain -0.7 dB · emolia-01120
Emotional Numbness ↓  /  Concentrationidentity +0.07 emotion 154 %   k-AB2-k4 · #10

This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Concentration up — by at least 0.25 each.

The chain starts with Concentration around average — 0.56, higher than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.34.

At the same time Emotional Numbness goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.06, then +0.09, then +0.19 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.79 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 39 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.716 before conversion and 0.781 after — it rose by 0.065. Neighbour-to-neighbour the worst pair went 0.807 → 0.865. (The earlier render, with segment 1 left raw, scores 0.626 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.344 in the original and +0.530 after conversion — 154 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, -0.299 became -0.280.

Quality. Mean predicted overall quality across the segments went 2.84 → 3.08 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.716 → 0.781 +0.065identity cos neighbours 0.807 → 0.865d_b rescored +0.344 → +0.530d_a rescored -0.299 → -0.280d_a mined -0.299d_b mined 0.343min_cos_consec (site) 0.7896min_cos_anchor (site) 0.7896dataset emolialang enspeaker EN_B00042_S04285total 38.4schain gain +3.0 dBseam step 1.9 dBcrossfades 150/150/150 ms
Script — 4 chunks, 4 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady
(emotional numbness, fear · monologue, authoritative) No (ahem) person in Warren's professional life who he detested more than Richard Nixon. (low mumble) Uhm, and seeing that that was about to happen or seeing that that could, (low mumble) uh, happen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, fear; style: monologue, authoritative; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.6/10; 8.0s, EN.
EN_B00042_S04285_W000225 · in -18.7 dBFS · gain -1.3 dB · emolia-01059
(malevolence malice, emotional numbness, triumph · monologue, casual) He (ahem) tried to resign, he submitted his resignation and made it contingent. He said he would leave the court upon the confirmation of his successor.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, emotional numbness, triumph; style: monologue, casual; good recording, no background noise; genuineness 3.2/6; vocal-burst blend 3.0/10; 7.5s, EN.
EN_B00042_S04285_W000226 · in -18.8 dBFS · gain -1.2 dB · emolia-01059
(monologue, conversational) But at the time, (ahem) uh, Johnson was a badly weakened president. (low mumble) Uh, he was not seeking re-election. The Vietnam War, uh, (low mumble) was upon the country. We're talking about the summer of 68.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, conversational; good recording, no background noise; genuineness 3.0/6; vocal-burst blend 3.8/10; 8.2s, EN.
EN_B00042_S04285_W000227 · in -19.8 dBFS · gain -0.2 dB · emolia-01059
(concentration, shame · formal, monologue) (ahem) Uh, it was, (low mumble) uh, too much of a stretch (low mumble) for John. Johnson then appointed Abe Fortas, uh, (low mumble) you know, who had been his personal lawyer and to whom he was very close. And the whole, uh, (ahem) confirmation really became a test of Johnson's ability to, to poll on the Senate. And the fact is he'd lost that ability by that point.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, shame; style: formal, monologue; average recording, no background noise; genuineness 2.3/6; vocal-burst blend 0.0/10; 15.3s, EN.
EN_B00042_S04285_W000228 · in -20.1 dBFS · gain +0.1 dB · emolia-01059
Concentration ↓  /  Fearidentity −0.03 emotion 22 %   k-AB2-k4 · #11

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Fear up — by at least 0.25 each.

The chain starts with Fear below average — 0.32, lower than 68 % of clips in this corpus — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.56.

At the same time Concentration goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.15, then +0.21, then +0.20 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 38 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.715 before conversion and 0.682 after — it fell by 0.033. Neighbour-to-neighbour the worst pair went 0.748 → 0.802. (The earlier render, with segment 1 left raw, scores 0.642 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.555 in the original and +0.123 after conversion — 22 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Concentration, -0.301 became -0.519.

Quality. Mean predicted overall quality across the segments went 2.96 → 3.19 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.715 → 0.682 -0.033identity cos neighbours 0.748 → 0.802d_b rescored +0.555 → +0.123d_a rescored -0.301 → -0.519d_a mined -0.303d_b mined 0.555min_cos_consec (site) 0.8511min_cos_anchor (site) 0.8833dataset emolialang zhspeaker ZH_B00002_S06693total 37.0schain gain +1.0 dBseam step 1.8 dBcrossfades 150/100/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, no disfluency, clear
(concentration · measured, steady, fairly narrow pitch, formal) 大家应观察增长率,随时间推移的加速和减速之间的相对增长率,以及每个组件与其自身历史百分比排名相比的增长率,这个表中存在一些重大分歧。首先。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: formal, monologue; average recording, no background noise; genuineness 0.0/6; vocal-burst blend 2.2/10; 15.6s, ZH.
ZH_B00002_S06693_W000009 · in -19.9 dBFS · gain -0.1 dB · emolia-03295
(normal-paced, steady, fairly narrow pitch, monologue) 如果你关注核心服务,不包括住房,美联储最密切关注的指标,你会发现这个指标远高于美联储百分之二点零的目标,而且是顽固的。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; average recording, no background noise; genuineness 0.4/6; vocal-burst blend 3.7/10; 10.9s, ZH.
ZH_B00002_S06693_W000010 · in -18.8 dBFS · gain -1.2 dB · emolia-03295
(relief · normal-paced, steady, fairly narrow pitch, formal) 在一个月和三个月的基础上,该指标的年化率分别为百分之四点九八和百分之五点一四。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief; style: formal, narration; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 2.8/10; 7.5s, ZH.
ZH_B00002_S06693_W000011 · in -19.1 dBFS · gain -0.8 dB · emolia-03295
(normal-paced, fairly steady, moderate pitch range, formal) 这些数字会让美联储感到非常不舒服。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; very good recording, no background noise; genuineness 0.1/6; vocal-burst blend 1.4/10; 3.5s, ZH.
ZH_B00002_S06693_W000012 · in -19.0 dBFS · gain -1.0 dB · emolia-03295
Concentration ↓  /  Affectionidentity +0.18 emotion 93 %   k-AB2-k4 · #12

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Affection up — by at least 0.25 each.

The chain starts with Affection clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.38.

At the same time Concentration goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.37. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.18, then +0.20 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.41 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.46 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.41, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 61 s · da · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.411 before conversion and 0.590 after — it rose by 0.179. Neighbour-to-neighbour the worst pair went 0.459 → 0.562. (The earlier render, with segment 1 left raw, scores 0.457 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Affection moved +0.380 in the original and +0.354 after conversion — 93 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.374 became -0.602.

Quality. Mean predicted overall quality across the segments went 3.05 → 3.39 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.411 → 0.590 +0.179identity cos neighbours 0.459 → 0.562d_b rescored +0.380 → +0.354d_a rescored -0.374 → -0.602d_a mined -0.374d_b mined 0.380min_cos_consec (site) 0.4571min_cos_anchor (site) 0.4128dataset eurospeechlang daspeaker denmark_20231M082_2024-04-total 60.2schain gain +1.7 dBseam step 1.6 dBcrossfades 150/100/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · average recording, quiet background
(concentration, intoxication altered states of consciousness · measured, subdued, slightly relaxed, monologue) (low mumble) yderligere i forhold til Arktis end det, der faktisk er tilfældet (low mumble) på nuværende tidspunkt, (low mumble) og at det jo i så tilfælde også uvægerlig ville kræve øget (low mumble) involvering af, dialog (low mumble) med og inddragelse af både Grønland og Færøerne.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, intoxication altered states of consciousness; style: monologue; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 2.2/10; 16.9s, DA.
denmark_20231M082_2024-04-19_0900_5594368_5611312 · in -22.4 dBFS · gain +2.4 dB · eurospeech-00512
(disgust · measured, subdued, slightly relaxed, monologue) Tak for det. Der er ikke flere korte bemærkninger, så vi siger tak til hr. Carsten Bach fra Liberal Alliance. Og jeg byder nu velkommen til fru Nanna W. Gotfredsen fra Moderaterne.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as disgust; style: monologue, didactic; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 0.0/10; 15.0s, DA.
denmark_20231M082_2024-04-19_0900_5611312_5626312 · in -24.6 dBFS · gain +4.6 dB · eurospeech-00512
(fatigue exhaustion, intoxication altered states of consciousness, confusion · measured, very low-energy, relaxed, whispered) Tak for ordet, formand. Og tak til statsministeren for redegørelsen og til hr. Flemming Møller Mortensen, som indledte her med at sige, (ahem)
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is slightly cool, slightly dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, slightly submissive, neutral openness; reads as fatigue exhaustion, intoxication altered states of consciousness, confusion; style: whispered, casual; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.8/10; 12.0s, DA.
denmark_20231M082_2024-04-19_0900_5636514_5648496 · in -23.9 dBFS · gain +3.9 dB · eurospeech-00512
(affection, confusion, intoxication altered states of consciousness · normal-paced, very low-energy, relaxed, casual) der er en hel masse i Folketinget, som har beskæftiget sig med (ahem) dette område i mange år, (low mumble) også intensivt. Jeg er ikke en af dem, (low mumble) jeg er ny her, og jeg er endnu nyere som (low mumble) grønlands- og
full caption & clip details
An elderly feminine voice; delivery is very low-energy, normal-paced, relaxed, moderately variable; timbre is neutral-toned, very dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, audible breath; affect is positive, slightly submissive, neutral openness; reads as affection, confusion, intoxication altered states of consciousness; style: casual, playful; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 5.3/10; 16.8s, DA.
denmark_20231M082_2024-04-19_0900_5648496_5665264 · in -22.3 dBFS · gain +2.3 dB · eurospeech-00512
Disgust ↓  /  Emotional Numbnessidentity −0.02 emotion 84 %   k-AB2-k4 · #13

This chain comes from the two-sided rule: it only counts if both emotions move — Disgust down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.26.

At the same time Disgust goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.14, then +0.02, then +0.10 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 39 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.886 before conversion and 0.871 after — it fell by 0.015. Neighbour-to-neighbour the worst pair went 0.894 → 0.863. (The earlier render, with segment 1 left raw, scores 0.749 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.262 in the original and +0.221 after conversion — 84 % of the delta retained, which is most of it. On the other named axis, Disgust, -0.284 became -0.280.

Quality. Mean predicted overall quality across the segments went 3.17 → 3.23 (+0.05) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.886 → 0.871 -0.015identity cos neighbours 0.894 → 0.863d_b rescored +0.262 → +0.221d_a rescored -0.284 → -0.280d_a mined -0.284d_b mined 0.263min_cos_consec (site) 0.9582min_cos_anchor (site) 0.9572dataset emolialang enspeaker EN_muKcn85o4uytotal 38.2schain gain +2.6 dBseam step 0.5 dBcrossfades 100/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(disgust · normal-paced, steady, newsreading, formal) This stem-based definition is equivalent to the more common definition of sauropsida, which Modesto and Anderson synonymized with reptilia, since the latter is better known and more frequently used
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust; style: newsreading, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 11.5s, EN.
EN_muKcn85o4uy_W000062 · in -16.7 dBFS · gain -3.3 dB · emolia-02416
(normal-paced, steady, newsreading, formal) Unlike most previous definitions of reptilia, however, Modesto and Anderson's definition includes birds, as they are within the clade that includes both lizards and crocodiles
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 10.4s, EN.
EN_muKcn85o4uy_W000063 · in -16.8 dBFS · gain -3.2 dB · emolia-02416
(normal-paced, fairly steady, formal, authoritative) == Taxonomy classification to order level of the reptiles, after Benton, 2014
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.8/10; 5.9s, EN.
EN_muKcn85o4uy_W000064 · in -16.7 dBFS · gain -3.3 dB · emolia-02416
(emotional numbness · measured, steady, formal, newsreading) The cladogram presented here illustrates the "'family tree' of reptiles, and follows a simplified version of the relationships found by M. S. Lee, in 2013
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.8s, EN.
EN_muKcn85o4uy_W000068 · in -16.2 dBFS · gain -3.8 dB · emolia-02416
Malevolence Malice ↓  /  Shameidentity −0.02 emotion REVERSED   k-AB2-k4 · #14

This chain comes from the two-sided rule: it only counts if both emotions move — Malevolence Malice down and Shame up — by at least 0.25 each.

The chain starts with Shame clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.27.

At the same time Malevolence Malice goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.47 (lower than 53 % of clips in this corpus), a change of -0.51. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.06, then +0.15, then +0.06 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.82 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 44 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.874 before conversion and 0.851 after — it fell by 0.023. Neighbour-to-neighbour the worst pair went 0.805 → 0.861. (The earlier render, with segment 1 left raw, scores 0.828 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The emotional move did not survive. Re-scored end to end, Shame moved +0.270 in the original and -0.060 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Malevolence Malice, -0.505 became -0.495.

Quality. Mean predicted overall quality across the segments went 2.97 → 3.22 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.874 → 0.851 -0.023identity cos neighbours 0.805 → 0.861d_b rescored +0.270 → -0.060d_a rescored -0.505 → -0.495d_a mined -0.505d_b mined 0.270min_cos_consec (site) 0.8170min_cos_anchor (site) 0.9081dataset emolialang zhspeaker ZH_B00039_S00060total 42.5schain gain +1.6 dBseam step 1.5 dBcrossfades 150/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, no background noise, normally alert, slightly relaxed
(malevolence malice, impatience and irritability, longing · normal-paced, moderate pitch range, monologue, didactic) 然而,我却四处碰壁,一个月下来,口袋里差不多瘾,空空如也。幸而一位在超级市场工作的朋友,把那里准备扔掉的过期食品偷偷接济,我才勉强度日。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, impatience and irritability, longing; style: monologue, didactic; average recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.1/10; 13.2s, ZH.
ZH_B00039_S00060_W000002 · in -19.8 dBFS · gain -0.2 dB · emolia-03662
(thankfulness gratitude, jealousy and envy, sadness · normal-paced, moderate pitch range, monologue, whispered) 最后我只剩下一美元,却怎么也舍不得把它花掉。因为上面满是我喜爱的歌星的亲笔签名。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as thankfulness gratitude, jealousy and envy, sadness; style: monologue, whispered; average recording, no background noise; genuineness 1.6/6; vocal-burst blend 1.5/10; 8.1s, ZH.
ZH_B00039_S00060_W000003 · in -20.5 dBFS · gain +0.5 dB · emolia-03662
(pain, thankfulness gratitude, longing · measured, moderate pitch range, whispered, monologue) 一天早晨,我在停车场留意到,一名男子坐在一辆破旧不堪的汽车里,一连两天,汽车都停在原地。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as pain, thankfulness gratitude, longing; style: whispered, monologue; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.2/10; 9.0s, ZH.
ZH_B00039_S00060_W000004 · in -20.4 dBFS · gain +0.3 dB · emolia-03662
(measured, fairly narrow pitch, whispered, monologue) 我心里纳闷,这么大的风雪,他待在那儿干什么?第三天早晨,当我走进那辆汽车时,那名男子把车窗摇下来,我停住脚步和他攀谈起来。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: whispered, monologue; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 0.3/10; 12.7s, ZH.
ZH_B00039_S00060_W000005 · in -21.9 dBFS · gain +1.9 dB · emolia-03662
Confusion ↓  /  Interestidentity −0.00 emotion 35 %   k-AB2-k4 · #15

This chain comes from the two-sided rule: it only counts if both emotions move — Confusion down and Interest up — by at least 0.25 each.

The chain starts with Interest clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.30.

At the same time Confusion goes the other way, from 0.92 (higher than 92 % of clips in this corpus) to 0.49 (lower than 51 % of clips in this corpus), a change of -0.43. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are -0.03, then +0.18, then +0.14 — not a clean run: step 1 moves back the other way by 0.03 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 57 s · da · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.953 before conversion and 0.952 after — it fell by 0.001. Neighbour-to-neighbour the worst pair went 0.956 → 0.920. (The earlier render, with segment 1 left raw, scores 0.872 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.299 in the original and +0.104 after conversion — 35 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Confusion, -0.432 became -0.759.

Quality. Mean predicted overall quality across the segments went 3.28 → 3.40 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.953 → 0.952 -0.001identity cos neighbours 0.956 → 0.920d_b rescored +0.299 → +0.104d_a rescored -0.432 → -0.759d_a mined -0.432d_b mined 0.298min_cos_consec (site) 0.9628min_cos_anchor (site) 0.9469dataset eurospeechlang daspeaker denmark_20161M020_2016-11-total 55.7schain gain +2.6 dBseam step 1.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult somewhat masculine voice · fairly smooth, quiet background
(confusion, anger, intoxication altered states of consciousness · fast, energised, neutral tension, casual) da den tidligere S-R-SF-regering ændrede reglerne. Det var efter den nuværende regerings opfattelse en klar fejl, at den tidligere regering lempede på reglerne.
full caption & clip details
A young adult somewhat masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, slightly guarded; reads as confusion, anger, intoxication altered states of consciousness; style: casual, dramatic; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 10.0/10; 16.3s, DA.
denmark_20161M020_2016-11-22_1300_17488271_17504576 · in -23.5 dBFS · gain +3.5 dB · eurospeech-00276
(sadness, longing, bitterness · normal-paced, normally alert, neutral tension, casual) Denne regering ønsker at skærpe reglerne på området og sende et klart og tydeligt signal til kriminelle udlændinge om, at vi i Danmark ikke vil give ophold til udlændinge, som bryder danske love.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as sadness, longing, bitterness; style: casual, playful; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 6.4/10; 11.5s, DA.
denmark_20161M020_2016-11-22_1300_17504576_17516080 · in -24.2 dBFS · gain +4.2 dB · eurospeech-00276
(intoxication altered states of consciousness · normal-paced, normally alert, slightly relaxed, didactic) Det fremgår således (ahem) direkte af lovforslaget, at udvisning kun kan undlades, hvis den med sikkerhed vil være i strid med Danmarks internationale forpligtelser. Flere kriminelle udlændinge skal udvises.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as intoxication altered states of consciousness; style: didactic, whispered; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 1.7/10; 15.2s, DA.
denmark_20161M020_2016-11-22_1300_17516080_17531264 · in -24.7 dBFS · gain +4.7 dB · eurospeech-00276
(interest · normal-paced, normally alert, slightly relaxed, didactic) Regeringens grundsynspunkt er, at Danmark skal leve op til sine internationale forpligtelser. Vi vil dog samtidig ikke lægge skjul på, at vi ønsker at udforske spillerummet inden for vores internationale forpligtelser.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest; style: didactic, authoritative; good recording, quiet background; genuineness 1.4/6; vocal-burst blend 1.0/10; 13.4s, DA.
denmark_20161M020_2016-11-22_1300_17531264_17544624 · in -24.2 dBFS · gain +4.2 dB · eurospeech-00276
Astonishment Surprise ↓  /  Contemplationidentity −0.00 emotion 99 %   k-AB2-k4 · #16

This chain comes from the two-sided rule: it only counts if both emotions move — Astonishment Surprise down and Contemplation up — by at least 0.25 each.

The chain starts with Contemplation around average — 0.51, right about the corpus median — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.39.

At the same time Astonishment Surprise goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.05, then +0.14 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 27 s · ko · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.803 before conversion and 0.802 after — it fell by 0.001. Neighbour-to-neighbour the worst pair went 0.800 → 0.723. (The earlier render, with segment 1 left raw, scores 0.714 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.388 in the original and +0.385 after conversion — 99 % of the delta retained, which is essentially all of it. On the other named axis, Astonishment Surprise, -0.364 became -0.474.

Quality. Mean predicted overall quality across the segments went 2.88 → 3.06 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.803 → 0.802 -0.001identity cos neighbours 0.800 → 0.723d_b rescored +0.388 → +0.385d_a rescored -0.364 → -0.474d_a mined -0.364d_b mined 0.388min_cos_consec (site) 0.8528min_cos_anchor (site) 0.8528dataset emolialang kospeaker KO_O5cjDEXTbXItotal 25.6schain gain +1.3 dBseam step 0.7 dBcrossfades 100/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(normal-paced, average clarity, authoritative, formal) 정비하는 곳에는 다이슨 성품기와 화장품들이 구비되어 있구요.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 2.9/10; 3.5s, KO.
KO_O5cjDEXTbXI_W000032 · in -19.5 dBFS · gain -0.5 dB · emolia-03110
(normal-paced, clear, formal, authoritative) 운동을 하러 왔는데 아무도 없어 이 넓은 피트니스를 저 혼자 1시간 정도 이용했습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 3.1/10; 5.0s, KO.
KO_O5cjDEXTbXI_W000033 · in -19.8 dBFS · gain -0.2 dB · emolia-03110
(normal-paced, clear, formal, authoritative) 운동을 마치고 간단한 샤워를 하고 수영장으로 향합니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 3.5/10; 3.7s, KO.
KO_O5cjDEXTbXI_W000034 · in -19.0 dBFS · gain -1.0 dB · emolia-03110
(measured, clear, monologue, narration) 사람이 없을 때 촬영하고 싶었지만 하필 제가 간 날, 다음 날이 한 달에 한 번 있는 수영장이 운영하지 않는 날이라 사람이 많은 시간에 촬영할 기회밖에 없었네요. 수영장 인테리어는 매우 고급지고,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.6/10; 13.9s, KO.
KO_O5cjDEXTbXI_W000035 · in -20.1 dBFS · gain +0.1 dB · emolia-03110
Embarrassment ↓  /  Interestidentity −0.03 emotion 109 %   k-AB2-k4 · #17

This chain comes from the two-sided rule: it only counts if both emotions move — Embarrassment down and Interest up — by at least 0.25 each.

The chain starts with Interest clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.27.

At the same time Embarrassment goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.13, then +0.07, then +0.06 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 60 s · el · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.916 before conversion and 0.890 after — it fell by 0.026. Neighbour-to-neighbour the worst pair went 0.928 → 0.874. (The earlier render, with segment 1 left raw, scores 0.814 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.266 in the original and +0.289 after conversion — 109 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Embarrassment, -0.252 became -0.122.

Quality. Mean predicted overall quality across the segments went 2.75 → 3.12 (+0.37) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.916 → 0.890 -0.026identity cos neighbours 0.928 → 0.874d_b rescored +0.266 → +0.289d_a rescored -0.252 → -0.122d_a mined -0.252d_b mined 0.266min_cos_consec (site) 0.9311min_cos_anchor (site) 0.9373dataset eurospeechlang elspeaker greece_olomeleia-20180208total 58.9schain gain +2.0 dBseam step 0.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: an adult feminine voice · fairly smooth, average recording, quiet background, moderately variable, some disfluency, wide pitch range
(embarrassment, confusion, doubt · brisk, normally alert, slightly relaxed, cartoonish) κινδύνεψαν να απολυθούν και αυτοί οι εργαζόμενοι, γιατί απολύθηκαν άλλοι εργαζόμενοι, αγαπητέ συνάδελφε. Βέβαια, πρέπει να σας πω ότι αυτό δεν έγινε ακριβώς γιατί σταμάτησε, πάγωσε, δεν υλοποιήθηκε αυτό, για να πούμε την αλήθεια.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as embarrassment, confusion, doubt; style: cartoonish, authoritative; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 2.0/10; 12.3s, EL.
greece_olomeleia-20180208_34719648_34731952 · in -20.9 dBFS · gain +0.9 dB · eurospeech-00751
(jealousy and envy, disgust, elation · brisk, energised, slightly relaxed, cartoonish) Το άλλο σημείο που πρέπει να πούμε εδώ είναι το δεύτερο σημείο, ότι το σύνολο των περιοχών «NATURA», (ahem) θαλασσίων και χερσαίων, θα καλυφθεί από φορείς διαχείρισης είτε μέσω της επέκτασης αρμοδιότητας των ήδη υφιστάμενων είτε με τη δημιουργία νέων φορέων διαχείρισης.
full caption & clip details
An adult feminine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as jealousy and envy, disgust, elation; style: cartoonish, playful; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 3.1/10; 16.1s, EL.
greece_olomeleia-20180208_34731952_34748080 · in -20.1 dBFS · gain +0.1 dB · eurospeech-00751
(sourness, contempt, shame · normal-paced, normally alert, neutral tension, cartoonish) Είναι σημαντικό να ειπωθεί ότι καμμία πλέον προστατευόμενη περιοχή δεν θα βρίσκεται εκτός φορέα διαχείρισης. Είναι δύο σημεία που είναι σημαντικά. (ahem) Αυτή η σημαντική εξέλιξη έχει πολλαπλά οφέλη. Γιατί;
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as sourness, contempt, shame; style: cartoonish, dramatic; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 4.3/10; 14.6s, EL.
greece_olomeleia-20180208_34748080_34762688 · in -19.3 dBFS · gain -0.7 dB · eurospeech-00751
(interest, malevolence malice, shame · brisk, energised, neutral tension, dramatic) ποικίλες πολιτικές που αφορούν στη βιώσιμη ανάπτυξη σε συνάρτηση με το φυσικό περιβάλλον και την τοπική κοινωνία μπορούν να εκπονηθούν από συνέργειες των φορέων διαχείρισης της τοπικής αυτοδιοίκησης και της κεντρικής εξουσίας. Έχουμε ένα τρίπτυχο, δηλαδή, πλέον (childlike giggle) δυνατότητας συνεργασίας και προοπτικής.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as interest, malevolence malice, shame; style: dramatic, cartoonish; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 4.3/10; 16.5s, EL.
greece_olomeleia-20180208_34762688_34779152 · in -19.2 dBFS · gain -0.8 dB · eurospeech-00751
Doubt ↓  /  Concentrationidentity +0.45 emotion 91 %   k-AB2-k4 · #18

This chain comes from the two-sided rule: it only counts if both emotions move — Doubt down and Concentration up — by at least 0.25 each.

The chain starts with Concentration clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.34.

At the same time Doubt goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.49 (lower than 51 % of clips in this corpus), a change of -0.47. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.10, then +0.21, then +0.03 — a plateau around step 3, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.38 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.38 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.38, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 63 s · da · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.377 before conversion and 0.830 after — it rose by 0.453. Neighbour-to-neighbour the worst pair went 0.377 → 0.830. (The earlier render, with segment 1 left raw, scores 0.693 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.337 in the original and +0.306 after conversion — 91 % of the delta retained, which is essentially all of it. On the other named axis, Doubt, -0.466 became -0.255.

Quality. Mean predicted overall quality across the segments went 3.07 → 3.32 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.377 → 0.830 +0.453identity cos neighbours 0.377 → 0.830d_b rescored +0.337 → +0.306d_a rescored -0.466 → -0.255d_a mined -0.466d_b mined 0.338min_cos_consec (site) 0.3837min_cos_anchor (site) 0.3837dataset eurospeechlang daspeaker denmark_20191M136_2020-06-total 61.8schain gain +1.0 dBseam step 1.0 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, balanced body, average recording, quiet background, some disfluency, light breath
(doubt, anger, confusion · normal-paced, normally alert, neutral tension, whispered) har de også fulgt den anmodning, der kom fra det daværende Indfødsretsudvalg, og hvis Indfødsretsudvalget som sagt ønsker en anden praksis, er det bare at skrive det.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, anger, confusion; style: whispered, casual; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 3.3/10; 11.0s, DA.
denmark_20191M136_2020-06-24_1300_14524288_14535264 · in -26.2 dBFS · gain +6.2 dB · eurospeech-00397
(triumph, embarrassment, shame · normal-paced, very low-energy, neutral tension, whispered) lang sagsbehandlingstid i forbindelse med statsborgerskabssagen og en helt anden form for (low mumble) situation, når man som Zohreh Bageri har søgt (low mumble) om opholdstilladelse
full caption & clip details
An adult masculine voice; delivery is very low-energy, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, embarrassment, shame; style: whispered, casual; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 4.9/10; 13.9s, DA.
denmark_20191M136_2020-06-24_1300_14535264_14549184 · in -26.8 dBFS · gain +6.8 dB · eurospeech-00397
(pride, triumph, concentration · normal-paced, normally alert, neutral tension, conversational) Hun er jo unægtelig, som folk, der søger opholdstilladelse, i en mere utryg situation, for de kan miste deres mulighed for at være i Danmark sammen med deres kære. Mener ministeren ikke, at det skaber utryghed hos borgere, hvis de vilkår, som de søger under, hele tiden ændres? Mener ministeren, det er rimeligt? Kl.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as pride, triumph, concentration; style: conversational, whispered; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 7.0/10; 18.1s, DA.
denmark_20191M136_2020-06-24_1300_14549184_14567328 · in -23.0 dBFS · gain +3.0 dB · eurospeech-00397
(concentration, contemplation, triumph · measured, subdued, slightly relaxed, monologue) Jeg mener faktisk, at det ville være fornuftigt, hvis udlændingelovgivningen, herunder reglerne om at få statsborgerskab, blev ændret noget sjældnere. Jeg vil gætte på, at udlændingeloven er en af de mest ændrede love, vi overhovedet har i Danmark. Den bliver hele tiden ændret – reglerne for familiesammenføring, for permanent ophold, (low mumble)
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, contemplation, triumph; style: monologue, whispered; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 5.5/10; 19.4s, DA.
denmark_20191M136_2020-06-24_1300_14567328_14586720 · in -24.3 dBFS · gain +4.3 dB · eurospeech-00397
Fear ↓  /  Disgustidentity +0.02 emotion REVERSED   k-AB2-k4 · #19

This chain comes from the two-sided rule: it only counts if both emotions move — Fear down and Disgust up — by at least 0.25 each.

The chain starts with Disgust below average — 0.33, lower than 67 % of clips in this corpus — and ends with it clearly present at 0.65, higher than 65 % of clips in this corpus. That is a total rise of 0.32.

At the same time Fear goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.58 (higher than 58 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.20, then +0.12 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 19 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.901 before conversion and 0.921 after — it rose by 0.020. Neighbour-to-neighbour the worst pair went 0.904 → 0.928. (The earlier render, with segment 1 left raw, scores 0.892 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disgust moved +0.202 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Fear, -0.307 became -0.145.

Quality. Mean predicted overall quality across the segments went 2.90 → 2.99 (+0.09) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.901 → 0.921 +0.020identity cos neighbours 0.904 → 0.928d_b rescored +0.202 → +0.000d_a rescored -0.307 → -0.145d_a mined -0.307d_b mined 0.322min_cos_consec (site) 0.9018min_cos_anchor (site) 0.9056dataset emolialang zhspeaker ZH_B00033_S01898total 17.9schain gain +2.2 dBseam step 0.9 dBcrossfades 150/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(authoritative, formal) 而菩提老祖交给孙悟空道家本领,法术中肯定也包括有医术在内。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 3.7/10; 5.8s, ZH.
ZH_B00033_S01898_W000018 · in -24.0 dBFS · gain +4.0 dB · emolia-03610
(authoritative, formal) 菩提老祖教的医术包罗万象,能够治百病,非常高级高深。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 3.6/10; 4.7s, ZH.
ZH_B00033_S01898_W000019 · in -22.5 dBFS · gain +2.5 dB · emolia-03610
(formal, authoritative) 不过因为猴子天性活泼好动,让唐僧一直对他不屑。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 2.0/10; 4.0s, ZH.
ZH_B00033_S01898_W000020 · in -22.9 dBFS · gain +2.9 dB · emolia-03610
(authoritative, formal) 猴头见了国王,一下惊吓了国王,于是提出悬思诊脉。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.9/10; 4.1s, ZH.
ZH_B00033_S01898_W000021 · in -22.9 dBFS · gain +2.9 dB · emolia-03610
Contemplation ↓  /  Contentmentidentity +0.03 emotion 108 %   k-AB2-k4 · #20

This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Contentment up — by at least 0.25 each.

The chain starts with Contentment clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.27.

At the same time Contemplation goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.11, then +0.05, then +0.12 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 39 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.756 before conversion and 0.782 after — it rose by 0.025. Neighbour-to-neighbour the worst pair went 0.711 → 0.710. (The earlier render, with segment 1 left raw, scores 0.734 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.274 in the original and +0.297 after conversion — 108 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contemplation, -0.285 became -0.393.

Quality. Mean predicted overall quality across the segments went 2.91 → 3.12 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.756 → 0.782 +0.025identity cos neighbours 0.711 → 0.710d_b rescored +0.274 → +0.297d_a rescored -0.285 → -0.393d_a mined -0.286d_b mined 0.274min_cos_consec (site) 0.8151min_cos_anchor (site) 0.8384dataset emolialang enspeaker EN_vUJi3ZDjt4ktotal 38.0schain gain +1.4 dBseam step 3.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · average recording, quiet background, normal-paced, normally alert, moderate pitch range
(contemplation, doubt, sadness · neutral tension, moderately variable, some disfluency, casual) You know, maybe, maybe you're, you're getting, getting ready to get kicked out of your house or something. Now that, that's a really hard situation. But remember, at least you still have your life. At least you still have, you know.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, dark, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as contemplation, doubt, sadness; style: casual, monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 5.9/10; 9.3s, EN.
EN_vUJi3ZDjt4k_W000036 · in -22.8 dBFS · gain +2.8 dB · emolia-00592
(infatuation, embarrassment, doubt · neutral tension, moderately variable, frequent disfluency, casual) (ahem) I don't know exactly your situation, but surely you understand what I'm saying. Find a reason to be happy. To be thankful. (ahem)
full caption & clip details
An elderly masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as infatuation, embarrassment, doubt; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 5.6/10; 9.5s, EN.
EN_vUJi3ZDjt4k_W000037 · in -21.5 dBFS · gain +1.5 dB · emolia-00592
(thankfulness gratitude, affection, longing · relaxed, fairly steady, frequent disfluency, casual) Thankfulness and contentment really go hand in hand if you really want to be happy. (ahem) Uhm, and, and when you, when you,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, slightly thin; slurred, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as thankfulness gratitude, affection, longing; style: casual, ASMR; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 4.5/10; 5.5s, EN.
EN_vUJi3ZDjt4k_W000038 · in -17.0 dBFS · gain -3.0 dB · emolia-00592
(contentment, hope enthusiasm optimism, elation · slightly relaxed, fairly steady, some disfluency, whispered) Take the time to be thankful every day. Just take the time. Stop your busy schedule and just focus in on being thankful and being content with what you have. (low mumble) Uhm, you will have joy and you will have peace. It's like Paul said, you know, with these things I will be content.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contentment, hope enthusiasm optimism, elation; style: whispered, casual; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 5.0/10; 14.3s, EN.
EN_vUJi3ZDjt4k_W000039 · in -22.1 dBFS · gain +2.1 dB · emolia-00592